Topic / Trend Rising

Compact and Sparse Open Models Expand Local Inference

Releases like Spark-X2.5, MiniCPM5-2B, K2-Horizon-MoVA and Qwen-Drive bring compact or sparse open models with lower active parameters and specialized capabilities. llama.cpp is adding support for new architectures such as Tencent Hy4, reinforcing local, edge and on-premise inference paths.

Detected: 2026-09-09 · Updated: 2026-09-09

Related Coverage

2026-09-08 LocalLLaMA

Qwen brings open weights to autonomous driving with Qwen-Drive-1.0-4B

Qwen has published an open-weight model derived from a 4B checkpoint and fine-tuned for driving. The full BF16 checkpoint is 9B. It signals Chinese labs pushing toward autonomous driving with open models, with implications for local inference, data s...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-07 LocalLLaMA

MiniCPM5-2B: OpenBMB Leads Open Weights Models Under 4B

OpenBMB has released MiniCPM5-2B, a 2-billion-parameter open weights model scoring 15 on the Artificial Analysis Intelligence Index v4.2, the highest among open models up to 4B. For self-hosted deployments, the small size and open weights lower the p...

#LLM On-Premise #Fine-Tuning #DevOps
2026-09-06 LocalLLaMA

Spark-X2.5 GGUF models gain llama.cpp support for local 1M-token inference

A llama.cpp pull request adds support for Spark-X2.5, two compact 4B and 1.7B parameter LLMs with a native context window of up to 1M tokens and a hybrid sliding-window attention architecture. The move lowers the barrier for local and self-hosted inf...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-04 LocalLLaMA

Tencent Hy4 lands on llama.cpp: a signal for local inference

Pull request #28127 adds support for the Tencent Hy4 preview architecture in llama.cpp. A concrete signal for teams evaluating on-premise stacks and data sovereignty, with no official benchmarks.

#Hardware #LLM On-Premise #DevOps
← Back to All Topics