llama.cpp speeds up Qwen3.8-Flash-Next: 55 tokens/s on four RTX 3090s
llama.cpp has merged support for Qwen3.8-Flash-Next and a community report shows 55 tokens/s with a Q4 GGUF on four RTX 3090s. An informal test that shifts attention from software to hardware: for local inference, multi-GPU cost and complexity remain...