Topic / Trend Rising

Self-Hosted LLM Inference Optimization

Local LLM deployments advance through quantization, MoE expert caching, KV cache optimization, and hardware-specific tuning across Apple Silicon, RTX, AMD and Intel Arc. New compact open-weight models and llama.cpp integrations make on-premise inference faster, cheaper and more capable.

Detected: 2026-09-08 · Updated: 2026-09-08

Related Coverage

2026-09-07 LocalLLaMA

llama.cpp, MCP and FreeCAD: a local LLM that designs geometry

A guide shows how to connect llama.cpp, the MCP protocol and FreeCAD to generate solids with a local model. The workflow uses Qwen3.8-27B quantized Q4_K_M and an mmproj-F16 vision projector: the model can call tools, read screenshots and verify geome...

#Hardware #LLM On-Premise
2026-09-06 LocalLLaMA

llama.cpp fork tests expert expansion on MoE models with Apple Metal

A custom llama.cpp branch brings expert expansion to MoE models and tests it on Apple Metal. The author says it beats their DS4 version, but cross-platform validation is missing. The case highlights how local inference for sparse models still depends...

#Hardware #LLM On-Premise #DevOps
2026-09-06 LocalLLaMA

DeepSeek-V4-Flash-Vision vs Qwen3.8-Flash-Next: slower tokens, faster tasks

A local comparison with Q8_K_XL quantization on two 128GB StrixHalo nodes shows that DeepSeek-V4-Flash-Vision generates tokens more slowly than Qwen3.8-Flash-Next, yet completes tasks in about half the time. The author observes fewer hallucinations a...

#Hardware #LLM On-Premise #DevOps
2026-09-06 LocalLLaMA

Spark-X2.5 GGUF models gain llama.cpp support for local 1M-token inference

A llama.cpp pull request adds support for Spark-X2.5, two compact 4B and 1.7B parameter LLMs with a native context window of up to 1M tokens and a hybrid sliding-window attention architecture. The move lowers the barrier for local and self-hosted inf...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-06 LocalLLaMA

Uncensored Qwen 3.8 27B: surgical edits beat aggressive model modifications

Eleven days and about 167 GPU hours to compare eight abliterated Qwen 3.8 27B variants against the base model. The data shows surgical edits beating aggressive ones: the most heavily edited models loop in their thinking and lose usability. Copyright ...

#Hardware #LLM On-Premise #DevOps
2026-09-06 LocalLLaMA

Qwen3.8-Flash-Next: 45 tok/s on M4 Max, 25 on M2 Ultra for local inference

The Qwen3.8-Flash-Next-oQ4e-mtp variant shows local inference speeds on Apple Silicon: 45 tokens per second on M4 Max and 25 on M2 Ultra, according to llm-bench.io. The source also notes similar speed to Qwen3.8 27B on Apple Silicon. Useful for self-...

#Hardware #LLM On-Premise #DevOps
2026-09-04 LocalLLaMA

Qwen3.8 27B on 16GB VRAM: benchmarking 21 quantized variants

A benchmark of 21 quantized Qwen3.8 27B variants on an RTX 5080 with 16GB VRAM shows how much GGUF choice matters. The bartowski IQ4_XS variant offers the lowest mean KL divergence and highest same-top-p agreement, while more aggressive quants prove ...

#Hardware #LLM On-Premise #DevOps
2026-09-04 LocalLLaMA

Tencent Hy4 lands on llama.cpp: a signal for local inference

Pull request #28127 adds support for the Tencent Hy4 preview architecture in llama.cpp. A concrete signal for teams evaluating on-premise stacks and data sovereignty, with no official benchmarks.

#Hardware #LLM On-Premise #DevOps
2026-09-04 LocalLLaMA

Qwen 3.8 Flash Next Locally Builds a Playable FPS

A user built a multiplayer FPS with a local Qwen 3.8 Flash Next model, Q4_K_XL quantization and 256k context, using opencode. Playable demo in two hours, three days of refinement, 20 tokens/s avg with MTP on RTX 5090 + RTX 4000 PRO. It highlights the...

#Hardware #LLM On-Premise #DevOps
2026-09-03 LocalLLaMA

Qwen-3.8-Next-Flash: Hot-Swappable Ngram Knowledge Injection in llama.cpp

An experimental modification to llama.cpp updates the Ngram PLE table of Qwen models in memory, injecting new knowledge without reloading the model. Tested only with q8 quantization, it requires memory mapping and suffers from unreliable output contr...

#LLM On-Premise #Fine-Tuning #DevOps
2026-09-03 LocalLLaMA

Arcon pushes Qwen3-4B: how much architecture matters beyond parameter count

The Arcon project uses Qwen3-4B with LoRA to build a local assistant with persistent memory, internal state, and tools. More than a chatbot, it is a test bed for how far a compact model can support assistant-like interaction. The answer matters for o...

#LLM On-Premise #Fine-Tuning #DevOps
2026-09-03 LocalLLaMA

Expert caching rewrites TCO: fluid MoE on aging on-premise hardware

The Qwen3.8-Flash-Next case on dual RTX 3090s shows that for MoE models VRAM is no longer the only variable: memory hierarchy and expert caching shift decode from 17 to 29 tokens/s. A concrete signal for teams evaluating on-premise deployments on old...

2026-09-03 LocalLLaMA

Why KV cache matters more than parameter count for local long-context models

For local LLMs, parameter count is not the whole constraint: the KV cache grows with every token and can become the real bottleneck at 100k or 200k context. GQA, MQA, and quantization reduce its footprint, but growth remains linear. Future optimizati...

#Hardware #LLM On-Premise #DevOps
2026-09-03 LocalLLaMA

Qwen3.8-Flash-Next on two RTX 3090s: expert cache pushes decode to 29 t/s

A user documents a jump from 17 to 25-29 tokens/s on two RTX 3090s with a full 261,888-token context, using llama.cpp's LRU expert cache and a smaller ubatch. The setup keeps 48 expert layers in system RAM and shows meaningful headroom for on-prem in...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-03 LocalLLaMA

Perplexity open-sources Mac inference server optimized for Qwen 3.6

Perplexity has open-sourced the lily repository inside pplx-garden, a Mac inference server built around a single model, Qwen 3.6, to get the best performance on Apple Silicon. This is not general-purpose serving: vertical integration around one model...

#Hardware #LLM On-Premise #DevOps
2026-09-03 Phoronix

Lemonade 11.9 Ships Experimental AMD ROCm HRX Backend via Llama.cpp

Lemonade 11.9, an AMD-backed open-source local AI server, arrives on Linux, Windows, and macOS with experimental ROCm HRX backend support for Llama.cpp. The project keeps its '100% free and private' focus on GPUs, CPUs, and NPUs. The most relevant te...

#Hardware #LLM On-Premise #DevOps
2026-09-02 Phoronix

Intel Brings vLLM 0.26 to Arc GPUs via Updated Docker Setup

Intel's LLM-Scaler-vLLM project releases a new beta based on vLLM 0.26, designed to get vLLM running on Arc and Arc Pro GPUs through Docker containers. The update lowers configuration friction for teams serving LLMs locally on Intel hardware, but it ...

#Hardware #LLM On-Premise #DevOps
← Back to All Topics