🏷️ LLM On-Premise

60 articles with this tag · all tags

Qwen3.8-27B beats Qwen3.6-27B in autonomous iteration on a BASIC ray tracer
LLM Qwen3.8-27B beats Qwen3.6-27B in autonomous iteration on a BASIC ray tracer 2026-08-16
Qwen3.8-27B runs locally and one-shots a Super Mario clone
Altro Qwen3.8-27B runs locally and one-shots a Super Mario clone 2026-08-15
RustConn 0.20 Polishes a GTK4 Connection Manager, Quietly Helping On-Prem Operations
Altro RustConn 0.20 Polishes a GTK4 Connection Manager, Quietly Helping On-Prem Operations 2026-08-15
KDE Plasma 6.8 adds fine-grained control over mouse and touchpad speed
Frameworks KDE Plasma 6.8 adds fine-grained control over mouse and touchpad speed 2026-08-15
Debian developers vote on LLM use in the project: governance and trust at stake
LLM Debian developers vote on LLM use in the project: governance and trust at stake 2026-08-15
Qwen 3.8 35BA3B appears in a commit: a signal before the launch
LLM Qwen 3.8 35BA3B appears in a commit: a signal before the launch 2026-08-15
Qwen 3.8 27B Release Day: Local Formats and the Deployment Shift
LLM Qwen 3.8 27B Release Day: Local Formats and the Deployment Shift 2026-08-15
Lemonade 11.6 Brings Muse-Glimmer 30B and Experimental ROCm Image Generation to Local AI
Frameworks Lemonade 11.6 Brings Muse-Glimmer 30B and Experimental ROCm Image Generation to Local AI 2026-08-14
LLM self-reflection: action routing beats diagnostic questions and taxonomies
LLM LLM self-reflection: action routing beats diagnostic questions and taxonomies 2026-08-14
Doom inside an LLM: a 34 GB checkpoint and token-based rendering
Hardware Doom inside an LLM: a 34 GB checkpoint and token-based rendering 2026-08-13
Writer targets token cost containment with new LLM based on GLM-5.2
LLM Writer targets token cost containment with new LLM based on GLM-5.2 2026-08-13
DeepSeek V4 Pro 0813 on Hugging Face: A Name Is Not Enough
LLM DeepSeek V4 Pro 0813 on Hugging Face: A Name Is Not Enough 2026-08-13
Qwen opens official countdown for Qwen3.8-27B on Hugging Face
LLM Qwen opens official countdown for Qwen3.8-27B on Hugging Face 2026-08-13
Nvidia doubles RTX PRO 6000 Blackwell price to $16,000, raising on-prem AI costs
Hardware Nvidia doubles RTX PRO 6000 Blackwell price to $16,000, raising on-prem AI costs 2026-08-13
Retrofitting Recurrent Depth into Pretrained LLMs: Faster Latent Reasoning with Sharp Limits
LLM Retrofitting Recurrent Depth into Pretrained LLMs: Faster Latent Reasoning with Sharp Limits 2026-08-13
Backtrader-Bench: Forcing LLMs to Run Code in Algorithmic Trading Benchmarks
Frameworks Backtrader-Bench: Forcing LLMs to Run Code in Algorithmic Trading Benchmarks 2026-08-13
AI Detectors Are Failing Academic Integrity by Penalizing Transparent Use
Altro AI Detectors Are Failing Academic Integrity by Penalizing Transparent Use 2026-08-13
Distribird brings Bayesian calibration to local, open-weight LLMs
Altro Distribird brings Bayesian calibration to local, open-weight LLMs 2026-08-13
Governing Conflicting LLMs: The Control Layer That Prevents Conversational Collapse
Frameworks Governing Conflicting LLMs: The Control Layer That Prevents Conversational Collapse 2026-08-13
Comma.ai launches Chestnut dock with AMD GPU and open-source firmware
Hardware Comma.ai launches Chestnut dock with AMD GPU and open-source firmware 2026-08-13
Supply-chain attack on LiteLLM exposes terabytes of credentials
Altro Supply-chain attack on LiteLLM exposes terabytes of credentials 2026-08-12
AI safety concerns grow, at Ai4 Hinton, Li and Ng discuss the value of openness
Altro AI safety concerns grow, at Ai4 Hinton, Li and Ng discuss the value of openness 2026-08-12
Linux Unlocks Hybrid Graphics on 2018–2019 MacBook Pros: A Boost for Local Inference, Too
Hardware Linux Unlocks Hybrid Graphics on 2018–2019 MacBook Pros: A Boost for Local Inference, Too 2026-08-12
Decoding the hidden reasoning of Claude and GPT: what it changes
Altro Decoding the hidden reasoning of Claude and GPT: what it changes 2026-08-12
LACT 0.10 Brings NVIDIA Overclocking and Blackwell Hotspot Sensing for Linux GPU Enthusiasts
Hardware LACT 0.10 Brings NVIDIA Overclocking and Blackwell Hotspot Sensing for Linux GPU Enthusiasts 2026-08-12
Intel LLM-Scaler Now Supports Muse Glimmer, Simplifying Local Inference on Arc (Pro) B GPUs
Frameworks Intel LLM-Scaler Now Supports Muse Glimmer, Simplifying Local Inference on Arc (Pro) B GPUs 2026-08-12
AMD: AI Agents Will Push the CPU-GPU Ratio Toward 1:1
Hardware AMD: AI Agents Will Push the CPU-GPU Ratio Toward 1:1 2026-08-12
EU Mandates Watermarking for Local Models Too — What It Means for Open Source AI
Altro EU Mandates Watermarking for Local Models Too — What It Means for Open Source AI 2026-08-12
Robust for conflict, weak for morality: LLM pipelines tested on French headlines
LLM Robust for conflict, weak for morality: LLM pipelines tested on French headlines 2026-08-12
LLM Agents Factory: An Agent Factory That Cuts Inference Costs
Frameworks LLM Agents Factory: An Agent Factory That Cuts Inference Costs 2026-08-12
More robust random neural networks: intuitionistic fuzzy takes on noisy data
Frameworks More robust random neural networks: intuitionistic fuzzy takes on noisy data 2026-08-12
How topology reveals the inner evolution of Transformers
LLM How topology reveals the inner evolution of Transformers 2026-08-12
Unsloth Desktop brings LLM training local: 2× faster, 70% less VRAM
Altro Unsloth Desktop brings LLM training local: 2× faster, 70% less VRAM 2026-08-11
Nemotron-3.5 Lightning: NVIDIA’s bet on efficiency for local inference
LLM Nemotron-3.5 Lightning: NVIDIA’s bet on efficiency for local inference 2026-08-11
FastFlowLM 1.0: AMD brings NPU AI under the ROCm umbrella
Frameworks FastFlowLM 1.0: AMD brings NPU AI under the ROCm umbrella 2026-08-11
Luth-2: French small language models beat 3x larger competitors, redefining local AI
LLM Luth-2: French small language models beat 3x larger competitors, redefining local AI 2026-08-11
LLMs and Waste Management: WuYuEval Reveals the Limits of Generalist AI
LLM LLMs and Waste Management: WuYuEval Reveals the Limits of Generalist AI 2026-08-11
Self-adaptive fuzzing exposes the hallucination cracks in multimodal LLMs
Frameworks Self-adaptive fuzzing exposes the hallucination cracks in multimodal LLMs 2026-08-11
Training a 1B LLM from scratch for under $200: the frontier of accessible self-hosting
LLM Training a 1B LLM from scratch for under $200: the frontier of accessible self-hosting 2026-08-11
Linux: Open-Source NVIDIA “Nova” Driver Gains Functionality with Rust
Hardware Linux: Open-Source NVIDIA “Nova” Driver Gains Functionality with Rust 2026-08-11
Fedora CoreOS enables systemd-oomd and zRAM swap: a signal for on-premise inference
Altro Fedora CoreOS enables systemd-oomd and zRAM swap: a signal for on-premise inference 2026-08-10
Ling-3.0-tiny: 8B parameters, 1.3B active, hitting 100 tokens/sec on MacBook
Altro Ling-3.0-tiny: 8B parameters, 1.3B active, hitting 100 tokens/sec on MacBook 2026-08-10
On-device agentic AI: Meta brings Muse Glimmer to ExecuTorch with speculative decoding and 128K context
Altro On-device agentic AI: Meta brings Muse Glimmer to ExecuTorch with speculative decoding and 128K context 2026-08-10
Muse-Glimmer-30B in GGUF: The Latest Piece of an Increasingly Mature Local Ecosystem
LLM Muse-Glimmer-30B in GGUF: The Latest Piece of an Increasingly Mature Local Ecosystem 2026-08-10
Meta Muse Glimmer: Local AI Agents on Consumer GPUs, 30B Parameter Model Released
LLM Meta Muse Glimmer: Local AI Agents on Consumer GPUs, 30B Parameter Model Released 2026-08-10
Meta's Muse Glimmer: A 30B Open-Weight Model Purpose-Built for Always-On Local Agents
LLM Meta's Muse Glimmer: A 30B Open-Weight Model Purpose-Built for Always-On Local Agents 2026-08-10
Linux 7.3 removes old SGI drivers: security and AI agent noise drive kernel pruning
Altro Linux 7.3 removes old SGI drivers: security and AI agent noise drive kernel pruning 2026-08-10
No cloud, just edge: Edgify raises $9M for AI that learns on devices
Altro No cloud, just edge: Edgify raises $9M for AI that learns on devices 2026-08-10
The inherited symlink and the long tail of technical debt: what Walter’s hack teaches us
Altro The inherited symlink and the long tail of technical debt: what Walter’s hack teaches us 2026-08-10
TEXAS Leverages Native MoE Routing for More Surgical Fine-Tuning
LLM TEXAS Leverages Native MoE Routing for More Surgical Fine-Tuning 2026-08-10
Truth is a vector: detecting fake news without leaving the model
LLM Truth is a vector: detecting fake news without leaving the model 2026-08-10
WeatherNext 2: DeepMind brings cyclone forecasting to a single H100 GPU
Hardware WeatherNext 2: DeepMind brings cyclone forecasting to a single H100 GPU 2026-08-09
BDH: Pathway's post-transformer architecture matches GPT-2 scaling on ordinary GPUs
LLM BDH: Pathway's post-transformer architecture matches GPT-2 scaling on ordinary GPUs 2026-08-09
Two flags take Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on a single DGX Spark
Frameworks Two flags take Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on a single DGX Spark 2026-08-09
Lophius: A Workbench for Transformer Research, Born from the Creator of Heretic
Frameworks Lophius: A Workbench for Transformer Research, Born from the Creator of Heretic 2026-08-09
Bare die and 3D-printed block push RTX 2060 Super to 28°C under load
Hardware Bare die and 3D-printed block push RTX 2060 Super to 28°C under load 2026-08-09
Meetily transcribes and summarizes meetings with no subscription, open source. Here’s how
Altro Meetily transcribes and summarizes meetings with no subscription, open source. Here’s how 2026-08-09
When AI writes drivers: DeepSeek generates a custom Metal kernel for Kimi K2 on Mac
Altro When AI writes drivers: DeepSeek generates a custom Metal kernel for Kimi K2 on Mac 2026-08-09
Kimi K3 slims to 478GB: the multilingual trim that shifts on-prem math
LLM Kimi K3 slims to 478GB: the multilingual trim that shifts on-prem math 2026-08-09
BitNet hits 36 tok/s on Xeon with a zero-dependency C engine—and crashes into the DRAM ceiling
Frameworks BitNet hits 36 tok/s on Xeon with a zero-dependency C engine—and crashes into the DRAM ceiling 2026-08-08