Show HN: Llama.cpp fork with 2-4x multiGPU speed for MoE models bigger than VRAM

(github.com)

1 points | by neuralll 6 hours ago ago

3 comments