Topic / Trend Rising

Small open models and data-efficient training

Compact open-weight LLMs under 4B-10B parameters and BabyLM results show gains from data organization, architecture and single-GPU training rather than parameter scale alone.

Detected: 2026-09-13 · Updated: 2026-09-13

Related Coverage

2026-09-13 LocalLLaMA

A dense 9.4B LLM ready for single-card training seeks a community

An independent researcher has opened a roughly 9.4B parameter dense model designed for single-card training. The project uses a Llama 3 tokenizer, distilled logit data, and architecture cues from Moonshot and Qwen. The real test is not technical; it ...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-11 ArXiv cs.CL

BabyLM 2026: A Principle-Driven Method Learns from 10 Million Words

Qiushi Engine ran an end-to-end autonomous research program on BabyLM 2026 Strict-Small, using 10 million corpus words and 100 million cumulative word presentations. Three stages linked frontier advancement, principle discovery, and principle-guided ...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-07 LocalLLaMA

MiniCPM5-2B: OpenBMB Leads Open Weights Models Under 4B

OpenBMB has released MiniCPM5-2B, a 2-billion-parameter open weights model scoring 15 on the Artificial Analysis Intelligence Index v4.2, the highest among open models up to 4B. For self-hosted deployments, the small size and open weights lower the p...

#LLM On-Premise #Fine-Tuning #DevOps
2026-09-06 LocalLLaMA

Spark-X2.5 GGUF models gain llama.cpp support for local 1M-token inference

A llama.cpp pull request adds support for Spark-X2.5, two compact 4B and 1.7B parameter LLMs with a native context window of up to 1M tokens and a hybrid sliding-window attention architecture. The move lowers the barrier for local and self-hosted inf...

#Hardware #LLM On-Premise #Fine-Tuning
← Back to All Topics