Topic / Trend Rising

Open-weights Self-hosted Ecosystem Expands

Beyond Qwen3.8, open-weights models and runtimes such as Ling-3.0, Ornith 1.5, GLM-Air, DeepSeek V4 Flash, and llama.cpp are broadening local serving options. Community reports test multi-GPU, CPU/GPU, and edge configurations for self-hosting.

Detected: 2026-08-26 · Updated: 2026-08-26

Related Coverage

2026-08-24 LocalLLaMA

DeepSeek V4 Flash in a basement: Epyc, RTX 5090 and 24 tokens/s

A user tested DeepSeek V4 Flash locally with an Epyc 7663, 256 GB of ECC RAM and an RTX 5090 32 GB. Using Q8_K_XL quantization and 100-128k token contexts, the system reached 23.8-24.6 tokens/s. The result shows an LLM with roughly 151 GB of weights ...

#Hardware #LLM On-Premise #DevOps
2026-08-24 LocalLLaMA

MTP lands in GLM-Air: 106B MoE speeds up on local GPUs

llama.cpp enables MTP for GLM-4.5-Air, a 106-billion-parameter MoE with only 12 billion active parameters. The change speeds up inference on memory-rich but compute-limited machines like Strix Halo, DGX Spark, and RTX 3090, and strengthens the fine-t...

#Hardware #LLM On-Premise #Fine-Tuning
2026-08-23 Tom's Hardware

RTX 5080 for $702 at Walmart: a signal for local LLM inference

A shopper found an RTX 5080 for $702 at Walmart, saving nearly $800 compared to current retail prices. More than a lucky break, the episode reveals the gap between official pricing and real-world cost of consumer GPUs, with direct effects on TCO for ...

#Hardware #LLM On-Premise #DevOps
2026-08-23 LocalLLaMA

Hosting Kimi K3 on 8 B300s: 92 tok/s and $190 per million tokens

A Modal hosting test with 8 B300s shows 92 tok/s in decode and $190 per million output tokens. The 1-bit variant on 8 A100s costs less per hour but triples the cost per token. The analysis reveals why hourly pricing is a misleading metric for LLM inf...

#Hardware #LLM On-Premise #DevOps
2026-08-22 Phoronix

Open-Source Etnaviv Driver Now Runs YOLOX

The open-source Etnaviv driver, originally created for Vivante GPUs and later extended to NPUs, can now run YOLOX. The development points to a practical alternative to proprietary SDKs for local edge inference.

#Hardware #LLM On-Premise #Fine-Tuning
2026-08-22 LocalLLaMA

Llama.cpp 0.2.0: local LLM runtime changes pace

The new llama.cpp release, with source and pre-built binaries, marks a step for self-hosted inference. Less friction for developers and companies running LLMs locally, more pressure on cloud-only services.

#Hardware #LLM On-Premise #DevOps
2026-08-20 LocalLLaMA

The boring path to DeepSeek V4 Flash at 140 tokens/s on 16 RTX 5060 Ti

A builder validated a self-hosted configuration with 16 RTX 5060 Ti 16GB cards, two Broadcom/PLX PEX88096 switches, and a Xeon Gold 6330. The system serves DeepSeek V4 Flash-0731 with up to 1 million tokens of context and, in one setup, an average ge...

#Hardware #LLM On-Premise #DevOps
2026-08-19 LocalLLaMA

Ornith 1.5: three models from 9B to 397B with GGUF versions for self-hosting

Three new Ornith 1.5 models—9B, 35B-A3B, and 397B—have appeared on Hugging Face, each with GGUF versions. The immediate availability of quantized formats signals a direct focus on local and self-hosted deployment, prompting reflection on the trade-of...

#Hardware #LLM On-Premise #DevOps
← Back to All Topics