LLM on-device: LFM2.5-2.6B runs at 17 tok/s on OnePlus 13, CPU-only inference
A developer ran LFM2.5-2.6B, a 2.69-billion-parameter LLM with a 128K context window, on a OnePlus 13 smartphone using CPU alone and Q4_K_M quantization. The custom inference engine, just 450 KB, hits 17 tok/s and supports other model architectures. ...