Vector Index Output Heads Speed Up CPU LLM Inference
Replacing the dense vocabulary projection with approximate HNSW-based search accelerates autoregressive decoding on CPU. Tested on Gemma 3, Llama 3.2 and Qwen 3, batch-one throughput improves by up to 82% for Gemma 3 270M while preserving AlpacaEval ...