Intel Battlemage and llama.cpp: A Patch Rewrites the TCO of On-Premise Inference
A handful of code lines in llama.cpp multiplies Intel Battlemage GPU performance up to 2.7 times on 118K-token contexts with quantized KV cache. The gain is so sharp that it turns a gaming card into a tangible alternative for local inference of 35B m...