Topic / Trend Rising

Local and On-Device LLM Inference Expands

A wave of guides, benchmarks and hardware releases shows LLMs running locally on Apple Silicon, RTX GPUs, edge SoCs and dedicated desktops. Market gray modifications and new on-device cores push VRAM, bandwidth and power envelopes for self-hosted inference.

Detected: 2026-09-12 · Updated: 2026-09-12

Related Coverage

2026-09-11 LocalLLaMA

China-modified 96 GB RTX 5090: gray market narrows the gap for local LLMs

A 96 GB VRAM RTX 5090 sold on Alibaba for under $4,000 shows growing gray-market demand for affordable on-premise inference. The Chinese mod triples the standard card's memory, lowering the barrier for running larger models locally. But without warra...

#Hardware #LLM On-Premise #DevOps
2026-09-10 ServeTheHome

Qualcomm and the Oryon, Adreno, Hexagon trio: on-device AI changes the rules

Qualcomm details the next-generation Oryon CPU, Adreno GPU, and Hexagon NPU. The move confirms that local inference becomes a primary design constraint for client chips and a strategic option for organizations evaluating self-hosted deployments and d...

#Hardware #LLM On-Premise #DevOps
2026-09-08 LocalLLaMA

GB per dollar and bandwidth: a compass for local LLM GPUs

A Reddit comparison uses VRAM per dollar and rated bandwidth to navigate GPUs discussed in LocalLLaMA communities. Prices were gathered with ChatGPT, new or second-hand, with acknowledged inaccuracies. A rough method, but useful for anyone evaluating...

#Hardware #LLM On-Premise #DevOps
2026-09-07 LocalLLaMA

llama.cpp, MCP and FreeCAD: a local LLM that designs geometry

A guide shows how to connect llama.cpp, the MCP protocol and FreeCAD to generate solids with a local model. The workflow uses Qwen3.8-27B quantized Q4_K_M and an mmproj-F16 vision projector: the model can call tools, read screenshots and verify geome...

#Hardware #LLM On-Premise
2026-09-06 LocalLLaMA

llama.cpp fork tests expert expansion on MoE models with Apple Metal

A custom llama.cpp branch brings expert expansion to MoE models and tests it on Apple Metal. The author says it beats their DS4 version, but cross-platform validation is missing. The case highlights how local inference for sparse models still depends...

#Hardware #LLM On-Premise #DevOps
2026-09-06 LocalLLaMA

DeepSeek-V4-Flash-Vision vs Qwen3.8-Flash-Next: slower tokens, faster tasks

A local comparison with Q8_K_XL quantization on two 128GB StrixHalo nodes shows that DeepSeek-V4-Flash-Vision generates tokens more slowly than Qwen3.8-Flash-Next, yet completes tasks in about half the time. The author observes fewer hallucinations a...

#Hardware #LLM On-Premise #DevOps
2026-09-06 LocalLLaMA

Qwen3.8-Flash-Next: 45 tok/s on M4 Max, 25 on M2 Ultra for local inference

The Qwen3.8-Flash-Next-oQ4e-mtp variant shows local inference speeds on Apple Silicon: 45 tokens per second on M4 Max and 25 on M2 Ultra, according to llm-bench.io. The source also notes similar speed to Qwen3.8 27B on Apple Silicon. Useful for self-...

#Hardware #LLM On-Premise #DevOps
← Back to All Topics