AI-Radar - Local LLMs, AI Hardware and Trends Observatory

AI-Radar for on-prem LLMs & Home AI

The daily radar on models, frameworks, and hardware to run AI locally. LLMs, LangChain, Chroma, mini-PCs, and everything you need for a distributed "in-house" brain.

⚙️ Stack: Local LLMs · LangChain · Transformers · ChromaDB · MiniPCs · AI boxes
🛰️ Ask Observatory (Q&A + RAG) connected to the article archive.
👥 160+ members · Join free →

⚡ Trending Now

View All →

🛠️ Guides & On-Premise Observatory

🚀 Run models locally → All guides →

Evergreen, hands-on references for running AI locally — hardware, cost, privacy and the full stack.

🖥️ LLM On-Premise Observatory Hardware, stack, governance and reference architectures for local AI.

Latest Analysis & Radar News

AI-generated articles from feeds, with space for human editorial layer above the raw content.

AlphaGenome Atlas: DeepMind mappa 9 miliardi di varianti del DNA
📁 Altro AI generated 🏆 IEEE Spectrum

Google DeepMind Maps 9 Billion Possible DNA Variants

DeepMind has released AlphaGenome Atlas, a public repository of precomputed predictions for nine billion possible single-letter changes to the human genome. It removes the need to write code or run the model, adds an impact score, and is available for non-commercial research. The dataset holds around one petabyte and required an 80-fold speedup, achieved through model distillation, GPU kernel optimization, and eliminating redundant calculations.

2026-09-08 📰 Source
Replit apre a Londra e sfida la monocultura della Silicon Valley
📁 Market AI generated ℹ️ Tech.eu

Replit Opens London Office and Challenges Silicon Valley's Monoculture

Replit has opened its first office outside the United States in London. The company currently employs around twenty people in the British capital and could reach one hundred by early next year. CEO Amjad Masad criticized Silicon Valley's monoculture, linking London's diversity to better idea generation. He also called for a level playing field and open source, accessible AI not controlled by a few companies.

2026-09-08 📰 Source
Arm lancia Total Design per la physical AI e un framework per la robotica
📁 Frameworks AI generated ℹ️ AI News

Arm launches Total Design for Physical AI and robotics framework

Arm brings together more than 80 partners to define common standards for robotics and physical AI. The Robotics Capability Framework introduces capability tiers similar to SAE levels, tying use cases to latency, compute placement, memory, power, and safety constraints. The initiative extends a method already used for cloud AI infrastructure to physical systems and signals a push toward local and edge compute.

2026-09-08 📰 Source
Cambricon nella PyTorch Foundation: il software è la vera partita dell'hardware
📁 Hardware AI generated ✅ PyTorch Blog

Cambricon joins PyTorch Foundation: software is the real hardware battle

Cambricon joining the PyTorch Foundation as a Platinum member signals a shift: controlling the open source software stack matters as much as chip performance. For teams evaluating self-hosted LLMs, native support for diverse backends reduces CUDA lock-in, maintenance burden, and TCO. The real test remains the continuity and maturity of its upstream contributions.

2026-09-08 📰 Source
Huawei Kirin 9050 Pro: LogicFolding porta densità +55% e consumi NPU -66%
📁 Hardware AI generated ✅ DigiTimes

Huawei Kirin 9050 Pro puts LogicFolding to the test with 55% density gain and 66% lower NPU power

Huawei has put LogicFolding to the test on the Kirin 9050 Pro: according to the source, the density gain reaches 55% and NPU power consumption drops by 66%. The data points to a clear direction for local inference: denser chips and more efficient NPUs change the calculations for those evaluating edge and on-premise deployments, reducing the energy required to run models on-device.

2026-09-08 📰 Source
Cambricon nella PyTorch Foundation: il peso dell'hardware open source
📁 Frameworks AI generated ✅ PyTorch Blog

Cambricon Joins PyTorch Foundation: The Weight of Open-Source Hardware

Cambricon joining the PyTorch Foundation as a Platinum member marks a strategic step for the open hardware ecosystem. The company, an AI chip pioneer since 2016, brings years of upstream contributions to PyTorch and vLLM, aiming for a more uniform device abstraction. For those evaluating on-premise deployments, this strengthens the option of self-hosted stacks on non-CUDA hardware, with implications for TCO and data sovereignty.

2026-09-08 📰 Source
PyTorch Foundation apre a Alibaba, Ant Group e Cambricon: la partita è il multi-backend
📁 Frameworks AI generated ✅ PyTorch Blog

PyTorch Foundation welcomes Alibaba, Ant Group and Cambricon: the multi-backend fight

Alibaba Cloud and Cambricon join the PyTorch Foundation as Platinum members, while Ant Group joins as Gold. In Shanghai, keynotes with Huawei outline an open stack spanning silicon, accelerators and agent runtimes. The real signal is a push to make PyTorch hardware-agnostic, with direct consequences for self-hosted infrastructure and alternatives to mainstream GPUs.

2026-09-08 📰 Source
MiniCPM5-2B: OpenBMB guida i modelli open weights sotto i 4B
📁 LLM AI generated ℹ️ LocalLLaMA

MiniCPM5-2B: OpenBMB Leads Open Weights Models Under 4B

OpenBMB has released MiniCPM5-2B, a 2-billion-parameter open weights model scoring 15 on the Artificial Analysis Intelligence Index v4.2, the highest among open models up to 4B. For self-hosted deployments, the small size and open weights lower the practical threshold for local inference and fine-tuning on proprietary data, while normal validation still applies.

2026-09-07 📰 Source
llama.cpp, MCP e FreeCAD: un LLM locale che progetta geometrie
📁 Altro AI generated ℹ️ LocalLLaMA

llama.cpp, MCP and FreeCAD: a local LLM that designs geometry

A guide shows how to connect llama.cpp, the MCP protocol and FreeCAD to generate solids with a local model. The workflow uses Qwen3.8-27B quantized Q4_K_M and an mmproj-F16 vision projector: the model can call tools, read screenshots and verify geometry. Everything stays on the machine. For on-prem CAD, a fully local agent-tool loop becomes a concrete scenario, with design files never leaving the company perimeter.

2026-09-07 📰 Source
A Bengaluru la sfida è passare da utenti a costruttori di sistemi ML
📁 Frameworks AI generated ✅ PyTorch Blog

In Bengaluru, the push from AI users to ML systems builders

More than 170 students, engineers, and researchers attended the Red Hat-Hugging Face event in Bengaluru. The focus was profiling, inference, RL environments, distributed training, and GPU communication. The message: India must move beyond AI consumption and enter the open source core stack.

2026-09-07 📰 Source
Rustls 0.23.44 attiva i certificati post-quantum ML-DSA di default
📁 Altro AI generated ✅ Phoronix

Rustls 0.23.44 enables post-quantum ML-DSA certificates by default

Rustls 0.23.44 enables ML-DSA certificates by default, moving post-quantum authentication from opt-in to standard behavior. For teams operating self-hosted TLS endpoints in front of LLM services, the shift reduces manual configuration but increases pressure on private CA tooling, HSMs, and legacy clients. It signals that post-quantum migration is entering the ordinary lifecycle of libraries.

2026-09-07 📰 Source
Astra e LLM open source: il divario vero è nell’orchestrazione
📁 OnPremise AI generated ℹ️ LocalLLaMA

Astra and open source LLMs: the real gap is orchestration

A Reddit post asks when open source LLMs will catch up with Astra. The signal is not about model weights, but about infrastructure: multimodal pipeline, inference runtime, memory and orchestration. For those evaluating on-premise stacks, the gap translates into trade-offs between control, latency and TCO. Open source can compete on specific tasks, but the integrated experience requires coordinated investments. AI-Radar analyses implications for local deployments and data sovereignty.

2026-09-07 📰 Source
Un benchmark su 25 pasti mostra che per le calorie non vince il modello più grande
📁 LLM AI generated ℹ️ LocalLLaMA

A 25-meal benchmark shows calorie estimation doesn't favor the biggest model

A test on 25 Nutrition5k meals compares seven multimodal LLMs on calorie estimation under a 20% error threshold. Open-weights Muse Spark 1.3 hits 48%, while Qwen 3.8 27B stops at 16%. The ranking doesn't track model size—a useful signal for anyone evaluating consumer hardware and hybrid deployment.

2026-09-07 📰 Source
Open source LLM e Astra: il divario che la community osserva
📁 LLM AI generated ℹ️ LocalLLaMA

Open source LLMs and Astra: the gap the community is watching

A Reddit question reopens the comparison between proprietary multimodal assistants and open source LLMs. Without benchmarks, what remains is a sense of an AI pace that is hard to follow. For those evaluating local stacks, the issue is not just the model, but integration, latency, and on-premise infrastructure.

2026-09-07 📰 Source
llama.cpp, una fork testa l'expert expansion sui modelli MoE su Apple Metal
📁 Frameworks AI generated ℹ️ LocalLLaMA

llama.cpp fork tests expert expansion on MoE models with Apple Metal

A custom llama.cpp branch brings expert expansion to MoE models and tests it on Apple Metal. The author says it beats their DS4 version, but cross-platform validation is missing. The case highlights how local inference for sparse models still depends on individual forks and limited testing.

2026-09-06 📰 Source
DeepSeek-V4-Flash-Vision vs Qwen3.8-Flash-Next: token lenti, task più rapidi
📁 LLM AI generated ℹ️ LocalLLaMA

DeepSeek-V4-Flash-Vision vs Qwen3.8-Flash-Next: slower tokens, faster tasks

A local comparison with Q8_K_XL quantization on two 128GB StrixHalo nodes shows that DeepSeek-V4-Flash-Vision generates tokens more slowly than Qwen3.8-Flash-Next, yet completes tasks in about half the time. The author observes fewer hallucinations and less overinterpretation, while Qwen stalls in aggressive modes. A reminder to evaluate on-premise models beyond raw generation speed.

2026-09-06 📰 Source
Spark-X2.5 in GGUF: llama.cpp apre la strada ai modelli compatti da 1M token
📁 LLM AI generated ℹ️ LocalLLaMA

Spark-X2.5 GGUF models gain llama.cpp support for local 1M-token inference

A llama.cpp pull request adds support for Spark-X2.5, two compact 4B and 1.7B parameter LLMs with a native context window of up to 1M tokens and a hybrid sliding-window attention architecture. The move lowers the barrier for local and self-hosted inference, while training on Huawei Ascend clusters signals an alternative hardware supply chain. The models target coding, agents, and local workflows, with implications for data sovereignty and TCO.

2026-09-06 📰 Source
GPT Astra insegna a Qwen Next a scolpire in Blender: coaching al posto del fine-tuning
📁 LLM AI generated ℹ️ LocalLLaMA

GPT Astra Teaches Qwen Next to Sculpt in Blender: Coaching Replaces Fine-Tuning

An alternative to distillation and fine-tuning uses Codex and GPT Astra to transfer operational procedures for 3D modeling to Qwen Next via OpenCode and MCP Blender. The open model executes tasks after being guided, while the frontier model burns Pro quota. A signal for local deployment and cost analysis.

2026-09-06 📰 Source
Qwen 3.8 27B senza censura: i tagli mirati battono le modifiche pesanti
📁 LLM AI generated ℹ️ LocalLLaMA

Uncensored Qwen 3.8 27B: surgical edits beat aggressive model modifications

Eleven days and about 167 GPU hours to compare eight abliterated Qwen 3.8 27B variants against the base model. The data shows surgical edits beating aggressive ones: the most heavily edited models loop in their thinking and lose usability. Copyright remains the hardest wall, while chemistry and biology become easy to unlock. Chat template forensics reveals hidden behavioral levers.

2026-09-06 📰 Source
Open source e laboratori di frontiera: quando la parità percepita sposta il mercato
📁 OnPremise AI generated ℹ️ LocalLLaMA

Open Source and Frontier Labs: When Perceived Parity Shifts the Market

Perceived parity between open source models and frontier labs, emerging in a critical domain, shifts competition from technical quality to cost and governance. Organizations handling sensitive data see self-hosted LLMs as a concrete option: TCO, latency, and control outweigh public benchmark scores. The signal points to token price pressure, demand for local inference hardware, and a fragile premium for closed APIs.

2026-09-06 📰 Source
TrueForge usa il 63% di token in meno dei managed agent a parità di task
📁 Frameworks AI generated ℹ️ LocalLLaMA

TrueForge uses 63% fewer tokens than managed agents on the same task set

A benchmark on 14 cross-system tasks and three MCP servers shows that TrueForge with Opus 4.8 solves 11/14 tasks like Claude Managed Agents, but uses 63% fewer tokens and costs 30% less per run. With GLM-5.2 the cost drops to $3. Native tracing, sandbox, and non-lossy context management are still missing.

2026-09-06 📰 Source
Open source e laboratori di frontiera: un divario ormai marginale
📁 Market AI generated ℹ️ LocalLLaMA

Open source and frontier labs: the gap is now marginal

A cybersecurity network builder reports no longer seeing meaningful differences between open source models and frontier lab releases. With the perceived technical gap shrinking, marketing pressures on token pricing intensify as companies eye public markets. Local models like DeepSeek V4 Flash are entering the comparison.

2026-09-05 📰 Source
AMD forma un team 'elite' per portare Rust nello stack GPU
📁 Altro AI generated ✅ Phoronix

AMD Forms 'Elite' Team to Push Rust Deep Into the GPU Stack

AMD is assembling a high-level developer team to bring Rust deeper into the GPU software stack, from firmware to drivers, shader compilers and other components. The goal is to reduce memory-related vulnerability classes and improve the reliability of the software controlling GPUs, with effects beyond a single vendor.

2026-09-05 📰 Source
Qwen3.8 27b su una 3090: coding agentico locale e il gap con i modelli grandi
📁 Hardware AI generated ℹ️ LocalLLaMA

Qwen3.8 27b on an RTX 3090: agentic coding goes local and reveals a gap with larger models

A technical post shows how Qwen3.8 27b in Q4_K_XL quantization runs on an RTX 3090 with 24 GB of VRAM and 100k token context. The author argues that small self-hosted LLMs capable of 80-90% of repetitive work erode cloud vendor revenue more than frontier models. A gap remains when moving to larger models: multi-GPU servers or DGX clusters are needed, and hardware cost opens a new TCO calculation.

2026-09-05 📰 Source
Tencent Hy4 su llama.cpp: il segnale per l'inference locale
📁 OnPremise AI generated ℹ️ LocalLLaMA

Tencent Hy4 on llama.cpp: local inference shifts its center of gravity

Pull request #28127, adding Hy4-preview support to llama.cpp, brings no benchmarks but signals convergence in the self-hosted ecosystem. Even cloud vendors are paying attention to on-premise pipelines. Stability, licensing, and documentation remain open questions: real value will be measured by operational costs, efficiency, and sovereignty constraints.

2026-09-05 📰 Source
Qwen3.8 27B su RTX 5080: 21 quantizzazioni alla prova dei 16 GB
📁 LLM AI generated ℹ️ LocalLLaMA

Qwen3.8 27B on 16GB VRAM: benchmarking 21 quantized variants

A benchmark of 21 quantized Qwen3.8 27B variants on an RTX 5080 with 16GB VRAM shows how much GGUF choice matters. The bartowski IQ4_XS variant offers the lowest mean KL divergence and highest same-top-p agreement, while more aggressive quants prove underwhelming. The data helps local inference users balance quality and VRAM capacity.

2026-09-04 📰 Source
768 GB di VRAM, 18.000 euro: il sogno self-hosted davanti ai modelli da 2 trilioni
📁 Hardware AI generated ℹ️ LocalLLaMA

768 GB of VRAM, €18k: the self-hosted dream hits the two-trillion-parameter wall

A builder spent €18,000 on an EPYC server with 768 GB of VRAM and 256 GB of RAM to run frontier open source LLMs locally. But upcoming models such as GLM 6 could be two to three times larger, making the machine unusable at 4-bit quantization. The case highlights the widening gap between self-hosted hardware and enterprise-scale requirements.

2026-09-04 📰 Source
Un FPS multiplayer nato da un LLM locale: Qwen 3.8 Flash Next in azione
📁 Altro AI generated ℹ️ LocalLLaMA

Qwen 3.8 Flash Next Locally Builds a Playable FPS

A user built a multiplayer FPS with a local Qwen 3.8 Flash Next model, Q4_K_XL quantization and 256k context, using opencode. Playable demo in two hours, three days of refinement, 20 tokens/s avg with MTP on RTX 5090 + RTX 4000 PRO. It highlights the growing maturity of self-hosted coding stacks and the hardware requirements for creative autonomy without cloud.

2026-09-04 📰 Source
Mesa 26.3 introduce il supporto alla modalità GPU a 64 bit di Intel Nova Lake P
📁 Hardware AI generated ✅ Phoronix

Mesa 26.3 Adds Support for Intel Nova Lake P's New 64-bit GPU Mode

Changes to Intel's graphics compiler merged into Mesa 26.3 reveal that Nova Lake P will adopt a 64-bit shader addressing mode. It is a deep architectural shift, not a minor update: it expands addressable memory and prepares the open-source driver for the next generation of Intel GPUs. For local stack evaluators, early support reduces the risk of hardware-software misalignment.

2026-09-04 📰 Source
Nvidia-Hugging Face: il controllo dell'AI passa da pagamenti e talenti
📁 OnPremise AI generated ✅ DigiTimes

Nvidia-Hugging Face: AI control moves through payments and talent

The Nvidia-Hugging Face agreement shows that control of the LLM ecosystem also moves through financial clauses. Paying silicon rivals and tying a billion dollars to personnel affects platform neutrality, talent costs, and dependencies for those evaluating self-hosted and on-premise stacks. A signal to monitor alongside VRAM, inference, and TCO.

2026-09-04 📰 Source
ASIC oltre le GPU: il segnale dal test per il 2027
📁 Hardware AI generated ✅ DigiTimes

ASICs overtaking GPUs by 2027: a signal from the test bench

Chunghwa Precision Test expects ASIC revenue to pass GPU revenue by 2027. The projection highlights a shift toward specialized AI chips, with consequences for TCO, workload flexibility and on-premises deployment. The real issue is the trade-off between per-token efficiency and the optionality needed to manage evolving models.

2026-09-04 📰 Source
Acer punta su RTX e GoogleBook mentre cresce la domanda di AI on-site
📁 Hardware AI generated ✅ DigiTimes

Acer bets on RTX and GoogleBook as on-site AI demand grows

Acer's decision to bet on RTX and GoogleBook signals a hardware shift toward local inference. On-site AI demand grows for sovereignty, latency, and cost control reasons, pushing vendors to invest in GPUs and devices for self-hosted workloads.

2026-09-04 📰 Source
sanoTTS: sintesi vocale completa su un microcontrollore da 3 dollari
📁 Altro AI generated ℹ️ LocalLLaMA

sanoTTS: a complete neural TTS stack on a $3 microcontroller

A new TTS stack with 294k parameters occupies 337 KB in INT8 and runs on a $3 microcontroller with 512 KB of SRAM and no NPU. It supports 11 voices and 6 languages; the 1.5M-parameter model beats larger networks on SCOREQ. A signal for local voice AI.

2026-09-03 📰 Source
Nvidia acquisisce Hugging Face: ora controlla la distribuzione dei modelli
📁 Market AI generated ℹ️ Tom's Hardware

Nvidia acquires Hugging Face: control over model distribution

With a $12.93 billion deal, Nvidia takes over Hugging Face, the go-to hub for model and dataset distribution. The move merges the leading GPU supplier with the channel that feeds fine-tuning and self-hosted inference. For teams running on-premise LLMs, incentives shift: platform neutrality becomes a critical point for sovereignty, TCO, and vendor dependency.

2026-09-03 📰 Source
NVIDIA acquisisce Hugging Face per 12,93 miliardi: neutralità sotto esame
📁 Market AI generated ℹ️ AI News

NVIDIA acquires Hugging Face for $12.93 billion: neutrality under scrutiny

NVIDIA has reached a $12.93 billion agreement to acquire Hugging Face. The deal promises to keep the platform open, multi-cloud, and independent of NVIDIA hardware. But controlling the main access point for open-weight LLMs raises questions about optimization, data sovereignty, and long-term incentives.

2026-09-03 📰 Source
Equinix spinge sull'AI distribuita con NVIDIA, Together AI, AWS e Google Cloud
📁 Altro AI generated ℹ️ TechWire Asia

Equinix pushes distributed AI with NVIDIA, Together AI, AWS and Google Cloud

Equinix introduced Inference Exchange, a distributed inference platform developed with NVIDIA and Together AI, and Fabric One, a managed connectivity service based on open specifications from AWS and Google Cloud. Inference Exchange lets enterprises place AI workloads closer to data and users, with multitenant and dedicated single-tenant options on reserved GPUs. Fabric One automates network provisioning through portals, APIs, agents, and natural-language prompts. Beta in 2026, general availability in 2027.

2026-09-03 📰 Source
NVIDIA acquisisce Hugging Face per 12,93 miliardi di dollari
📁 Market AI generated ✅ Phoronix

NVIDIA Acquires Hugging Face for $12.93 Billion

NVIDIA has announced the acquisition of Hugging Face for $12.93 billion. The deal combines the leading GPU supplier for LLM workloads with a central hub for models, datasets and developer tools, with broad consequences for self-hosted and on-premise AI deployments.

2026-09-03 📰 Source
Nvidia-Hugging Face: la voce da 12,9 miliardi che scuote il self-hosted
📁 Market AI generated ℹ️ LocalLLaMA

Nvidia-Hugging Face: the $12.9 billion rumor shaking the self-hosted AI stack

A Reddit post claims Nvidia plans to acquire Hugging Face for $12.9 billion. The report remains unconfirmed, but if true it would reshape the neutral layer between GPU hardware and open-weight LLM distribution. For on-premise and self-hosted deployments, the key question is dependence on a single vendor across the entire stack.

2026-09-03 📰 Source
Qwen-3.8-Next-Flash: iniezione di conoscenza a caldo in llama.cpp
📁 Frameworks AI generated ℹ️ LocalLLaMA

Qwen-3.8-Next-Flash: Hot-Swappable Ngram Knowledge Injection in llama.cpp

An experimental modification to llama.cpp updates the Ngram PLE table of Qwen models in memory, injecting new knowledge without reloading the model. Tested only with q8 quantization, it requires memory mapping and suffers from unreliable output control. It points toward hot-swappable memory for self-hosted LLMs, but memory and predictability constraints remain.

2026-09-03 📰 Source
Arcon spinge Qwen3-4B: quanto conta l'architettura, non solo i parametri
📁 Altro AI generated ℹ️ LocalLLaMA

Arcon pushes Qwen3-4B: how much architecture matters beyond parameter count

The Arcon project uses Qwen3-4B with LoRA to build a local assistant with persistent memory, internal state, and tools. More than a chatbot, it is a test bed for how far a compact model can support assistant-like interaction. The answer matters for on-premise deployments and data sovereignty.

2026-09-03 📰 Source
La cache degli esperti riscrive il TCO: MoE fluido su hardware on-premise datato
📁 OnPremise AI generated ℹ️ LocalLLaMA

Expert caching rewrites TCO: fluid MoE on aging on-premise hardware

The Qwen3.8-Flash-Next case on dual RTX 3090s shows that for MoE models VRAM is no longer the only variable: memory hierarchy and expert caching shift decode from 17 to 29 tokens/s. A concrete signal for teams evaluating on-premise deployments on older hardware.

2026-09-03 📰 Source
Il calcolo lento come scelta: un Qwen 27B e la produttività asincrona
📁 Altro AI generated ℹ️ LocalLLaMA

Slow compute as a choice: a Qwen 27B and asynchronous productivity

A Reddit post reports a drop from 15 hours of active programming to 4 hours for debugging or implementing a feature with Qwen 27B, even at 0.5 tokens per second. The data reframes expectations for asynchronous inference and the value of recovered human time.

2026-09-03 📰 Source
La KV cache pesa più dei parametri sui modelli locali a contesto lungo
📁 LLM AI generated ℹ️ LocalLLaMA

Why KV cache matters more than parameter count for local long-context models

For local LLMs, parameter count is not the whole constraint: the KV cache grows with every token and can become the real bottleneck at 100k or 200k context. GQA, MQA, and quantization reduce its footprint, but growth remains linear. Future optimization may focus more on persistent memory and data movement than on parameters.

2026-09-03 📰 Source
Qwen3.8-Flash-Next su due RTX 3090: la cache esperti spinge il decode a 29 t/s
📁 Hardware AI generated ℹ️ LocalLLaMA

Qwen3.8-Flash-Next on two RTX 3090s: expert cache pushes decode to 29 t/s

A user documents a jump from 17 to 25-29 tokens/s on two RTX 3090s with a full 261,888-token context, using llama.cpp's LRU expert cache and a smaller ubatch. The setup keeps 48 expert layers in system RAM and shows meaningful headroom for on-prem inference on older hardware, provided the cache and buffers are sized to the available VRAM.

2026-09-03 📰 Source
Quando tacere è la risposta giusta: addestrare il confine dell'evidenza
📁 LLM AI generated 🏆 ArXiv cs.CL

When silence is the right answer: training the evidence boundary

A training framework teaches grounded QA models to answer only when evidence becomes sufficient, locating the exact point where abstention gives way to a response. On HotpotQA, 2WikiMultiHopQA and MuSiQue, with Qwen2.5-3B-Instruct and LoRA, the method lowers unsupported-answer rates on external sets.

2026-09-03 📰 Source
PRO-Step corregge il RAG multi-hop: premia ogni passo, non solo il risultato
📁 LLM AI generated 🏆 ArXiv cs.CL

PRO-Step: Step-Level Rewards for More Reliable Multi-Hop RAG

PRO-Step introduces step-level supervision for Retrieval-Augmented Generation. Instead of evaluating only the final answer, a generative process reward model checks logical validity and evidential grounding at every step of multi-hop reasoning. The method uses tree search and step-level direct preference optimization. Tests show best average exact match and F1 across five benchmarks. It signals a useful shift for more reliable self-hosted pipelines.

2026-09-03 📰 Source
Perplexity apre il server di inference per Mac ottimizzato per Qwen 3.6
📁 Frameworks AI generated ℹ️ LocalLLaMA

Perplexity open-sources Mac inference server optimized for Qwen 3.6

Perplexity has open-sourced the lily repository inside pplx-garden, a Mac inference server built around a single model, Qwen 3.6, to get the best performance on Apple Silicon. This is not general-purpose serving: vertical integration around one model changes the trade-offs among flexibility, maintenance, and performance. The piece explores who benefits and what it signals for local deployments.

2026-09-03 📰 Source
Lemonade 11.9 porta il backend sperimentale AMD ROCm HRX in Llama.cpp
📁 Frameworks AI generated ✅ Phoronix

Lemonade 11.9 Ships Experimental AMD ROCm HRX Backend via Llama.cpp

Lemonade 11.9, an AMD-backed open-source local AI server, arrives on Linux, Windows, and macOS with experimental ROCm HRX backend support for Llama.cpp. The project keeps its '100% free and private' focus on GPUs, CPUs, and NPUs. The most relevant technical detail is the early HRX integration, signaling another step toward local alternatives to CUDA for self-hosted deployments.

2026-09-03 📰 Source
GNOME Mutter 51.rc porta lo scan-out diretto dei buffer tra GPU diverse
📁 Altro AI generated ✅ Phoronix

GNOME Mutter 51.rc Adds Cross-GPU Buffer Scan-Out Support

The GNOME Mutter 51 release candidate introduces cross-GPU buffer scan-out. It lets the compositor avoid an intermediate copy when rendering and display happen on different GPUs—small but meaningful for Linux multi-GPU workstations.

2026-09-02 📰 Source
Greg KH: Linux 7.3 sarà un ciclo «rough» per il rumore dei contributi AI
📁 Altro AI generated ✅ Phoronix

Greg KH warns Linux 7.3 kernel cycle will be rough amid AI contribution noise

The first release candidate of Linux 7.3 has just arrived and Greg Kroah-Hartman expects a rough cycle due to AI and LLM noise. The rise in assisted bug reports and patches shifts costs onto maintainers. For those managing self-hosted infrastructure, kernel stability becomes a critical variable, affecting regressions, review times and long-term support choices.

2026-09-02 📰 Source
Intel aggiorna LLM-Scaler-vLLM: vLLM 0.26 su GPU Arc con Docker
📁 Frameworks AI generated ✅ Phoronix

Intel Brings vLLM 0.26 to Arc GPUs via Updated Docker Setup

Intel's LLM-Scaler-vLLM project releases a new beta based on vLLM 0.26, designed to get vLLM running on Arc and Arc Pro GPUs through Docker containers. The update lowers configuration friction for teams serving LLMs locally on Intel hardware, but it remains a beta: a sign of a maturing software ecosystem rather than a production-ready solution.

2026-09-02 📰 Source
← Previous Page 1 / 63 Next →
View Full Archive 🗄️

AI-Radar is an independent observatory covering AI models, local LLMs, on-premise deployments, hardware, and emerging trends. We provide daily analysis and editorial coverage for developers, engineers, and organizations exploring local AI solutions.

AI-RADAR badge LaunchTry LAUNCHING SOON ON LaunchTry Fazier badge