🏷️ RAG

60 articles with this tag · all tags

NVIDIA Nemotron Parse 2.0 Brings Chart-Aware Document Parsing to Self-Hosted Pipelines
LLM NVIDIA Nemotron Parse 2.0 Brings Chart-Aware Document Parsing to Self-Hosted Pipelines 2026-08-06
Xberg v1 is a Rust framework for truly local document intelligence
Frameworks Xberg v1 is a Rust framework for truly local document intelligence 2026-08-02
From a 5090 to a Mini Datacenter: The Parable of Trying to Escape API Fees
Altro From a 5090 to a Mini Datacenter: The Parable of Trying to Escape API Fees 2026-07-30
LLM Memory: A Wiki Pattern to Remember Dead Ends
Frameworks LLM Memory: A Wiki Pattern to Remember Dead Ends 2026-07-29
Small LLMs: the real on-premise frontier — tech community digs into practical uses
LLM Small LLMs: the real on-premise frontier — tech community digs into practical uses 2026-07-25
Solar-Open2: A 15B-Active MoE Model Targeting Agentic Workloads On-Premise
LLM Solar-Open2: A 15B-Active MoE Model Targeting Agentic Workloads On-Premise 2026-07-22
RIMS: Soft Aggregation Makes Small LLMs More Precise in Noisy RAG
LLM RIMS: Soft Aggregation Makes Small LLMs More Precise in Noisy RAG 2026-07-21
Microsoft to block screenshots of sensitive PDFs — but only in Edge
Altro Microsoft to block screenshots of sensitive PDFs — but only in Edge 2026-07-20
Byte-Exact KV Cache Grafting on Gemma 4 Turns Verified Knowledge into a Reusable State
LLM Byte-Exact KV Cache Grafting on Gemma 4 Turns Verified Knowledge into a Reusable State 2026-07-19
Bonsai 27B on iPhone: 27B LLM in 3.9GB with 1-bit quantization
LLM Bonsai 27B on iPhone: 27B LLM in 3.9GB with 1-bit quantization 2026-07-17
HG-RAG: Retrieval That Climbs Knowledge Trees for a Sharper LLM
Frameworks HG-RAG: Retrieval That Climbs Knowledge Trees for a Sharper LLM 2026-07-17
TSMC ramps up US fabs as Taiwan designers hesitate
Hardware TSMC ramps up US fabs as Taiwan designers hesitate 2026-07-17
France's Luciole-23B Lights an Open LLM Path for On-Premises AI
LLM France's Luciole-23B Lights an Open LLM Path for On-Premises AI 2026-07-16
CANDI-QA Exposes LLM Limits in Critical Sectors: Symbolic Reasoning Needed
LLM CANDI-QA Exposes LLM Limits in Critical Sectors: Symbolic Reasoning Needed 2026-07-15
Meta’s Adam Mosseri says AI token budgets could soon be capped per engineer
Market Meta’s Adam Mosseri says AI token budgets could soon be capped per engineer 2026-07-14
Bilibili’s Index-1.9B: A Small Open LLM That Challenges Much Larger Models
LLM Bilibili’s Index-1.9B: A Small Open LLM That Challenges Much Larger Models 2026-07-14
AuditWeave: A Tamper-Evident Ledger for AI Decision Tracking
Altro AuditWeave: A Tamper-Evident Ledger for AI Decision Tracking 2026-07-14
AI demand quadruples Taiwan memory revenue: a warning bell for on-premise deployments
Market AI demand quadruples Taiwan memory revenue: a warning bell for on-premise deployments 2026-07-14
Cloudflare’s AI crawler block: why on-premise AI deployments must rethink data costs
Altro Cloudflare’s AI crawler block: why on-premise AI deployments must rethink data costs 2026-07-13
Linux 7.2 Enables UltraRISC by Default: A Building Block for Open AI Hardware
Hardware Linux 7.2 Enables UltraRISC by Default: A Building Block for Open AI Hardware 2026-07-12
The End of Reactive Agents: Context Graph Brings Enterprise AI That Speaks Before You Ask
Frameworks The End of Reactive Agents: Context Graph Brings Enterprise AI That Speaks Before You Ask 2026-07-10
If You Already Pay for Cloud LLM, Local Embeddings and Rerankers Outweigh Local Language Models
Altro If You Already Pay for Cloud LLM, Local Embeddings and Rerankers Outweigh Local Language Models 2026-07-09
Perplexity is quietly building an AI coding tool to rival Cursor and Claude Code
Altro Perplexity is quietly building an AI coding tool to rival Cursor and Claude Code 2026-07-08
Local LLMs, no accuracy without RAG: a developer's benchmark
Altro Local LLMs, no accuracy without RAG: a developer's benchmark 2026-07-08
Prompt injection: how mass-market AI tools become arsenals for botnets
Altro Prompt injection: how mass-market AI tools become arsenals for botnets 2026-07-08
Longsys sees explosive profit growth in 1H26 fueled by AI memory demand
Market Longsys sees explosive profit growth in 1H26 fueled by AI memory demand 2026-07-06
Phison: AI demand could end NAND memory's boom-and-bust cycle
Hardware Phison: AI demand could end NAND memory's boom-and-bust cycle 2026-07-06
A 270 million parameter LLM built from scratch and designed for local inference
LLM A 270 million parameter LLM built from scratch and designed for local inference 2026-07-05
Local LLMs and agentic workloads: prefill is everything, KV heads beat parameters
LLM Local LLMs and agentic workloads: prefill is everything, KV heads beat parameters 2026-07-05
Fine-tuned Gemma 4 31B for copywriting: +290 Elo and no more clichés
LLM Fine-tuned Gemma 4 31B for copywriting: +290 Elo and no more clichés 2026-07-02
Intel publishes initial GCC patches for AI Compute Extensions (ACE)
Hardware Intel publishes initial GCC patches for AI Compute Extensions (ACE) 2026-07-02
SK Hynix bets $51bn on NAND plant to catch the AI memory wave
Hardware SK Hynix bets $51bn on NAND plant to catch the AI memory wave 2026-07-02
Memora: The Scalable Memory for AI Agents That Reduces Tokens by 98%
Frameworks Memora: The Scalable Memory for AI Agents That Reduces Tokens by 98% 2026-06-29
NASA tests local LLM inference for medical AI in deep space
Altro NASA tests local LLM inference for medical AI in deep space 2026-06-29
Local NPC Engine with Lightweight LLMs: The On-Premise Bet for Future RPGs
Altro Local NPC Engine with Lightweight LLMs: The On-Premise Bet for Future RPGs 2026-06-28
The next AI won’t be powered by better models alone
Altro The next AI won’t be powered by better models alone 2026-06-27
On-prem LLMs: the workflow you wish you had discovered sooner
Altro On-prem LLMs: the workflow you wish you had discovered sooner 2026-06-26
When you don’t have a data center GPU: strategies for local LLMs without a supercomputer
Hardware When you don’t have a data center GPU: strategies for local LLMs without a supercomputer 2026-06-26
Scaled Cognition raises $100M to build hallucination-free AI
LLM Scaled Cognition raises $100M to build hallucination-free AI 2026-06-25
Ornith-1.0: New LLM Family on Hugging Face, from 9B Dense to 397B MoE
LLM Ornith-1.0: New LLM Family on Hugging Face, from 9B Dense to 397B MoE 2026-06-25
NAND shortage until 2027: Phison warns as orders are booked through Q2
Hardware NAND shortage until 2027: Phison warns as orders are booked through Q2 2026-06-25
Linux 7.2: MGLRU improvement pushes MongoDB throughput up to 100% higher
Altro Linux 7.2: MGLRU improvement pushes MongoDB throughput up to 100% higher 2026-06-24
AMD Exposes Gamma 2.4 and 2.6 Curves in Linux Driver, Boosting Color Precision
Hardware AMD Exposes Gamma 2.4 and 2.6 Curves in Linux Driver, Boosting Color Precision 2026-06-24
The web data layer: AI’s new infrastructure frontier
Altro The web data layer: AI’s new infrastructure frontier 2026-06-24
Compri raises €3.2M for AI procurement, putting on-premise deployment in focus
Market Compri raises €3.2M for AI procurement, putting on-premise deployment in focus 2026-06-24
China’s LineShine tops TOP500: a HPC crown that doesn’t mean AI supremacy
Altro China’s LineShine tops TOP500: a HPC crown that doesn’t mean AI supremacy 2026-06-24
Baden Bower’s AI Visibility Index: How AI is rewriting the rules of digital visibility
Market Baden Bower’s AI Visibility Index: How AI is rewriting the rules of digital visibility 2026-06-22
IEEE Launches Virtual Training Course to Master Large Language Models
LLM IEEE Launches Virtual Training Course to Master Large Language Models 2026-06-19
Google AI mistakes horror fan-fiction for real-world facts
LLM Google AI mistakes horror fan-fiction for real-world facts 2026-06-19
Liquid AI releases two multilingual embedding models optimized for local retrieval
LLM Liquid AI releases two multilingual embedding models optimized for local retrieval 2026-06-18
Idle Multi-GPU Node? How to Repurpose Aging Hardware for Local LLM Inference
Hardware Idle Multi-GPU Node? How to Repurpose Aging Hardware for Local LLM Inference 2026-06-18
LLM Distillation: The Compute Challenge for GLM 5.2 Datasets
LLM LLM Distillation: The Compute Challenge for GLM 5.2 Datasets 2026-06-18
SproutRAG: Hierarchical RAG and Attention for Efficient Long-Document Management
Frameworks SproutRAG: Hierarchical RAG and Attention for Efficient Long-Document Management 2026-06-18
CaVe-VLM-CoT: An Interpretable Framework for Reliable Vision-Language Models
Frameworks CaVe-VLM-CoT: An Interpretable Framework for Reliable Vision-Language Models 2026-06-18
Extend.ai Releases Open Source UI Kit for Document Agents and RAG
Frameworks Extend.ai Releases Open Source UI Kit for Document Agents and RAG 2026-06-17
The Rise of Local Large Language Models: From "Toys" to Essential Tools
LLM The Rise of Local Large Language Models: From "Toys" to Essential Tools 2026-06-17
KPMG Withdraws AI Report: 'Hallucinations' Question Reliability
LLM KPMG Withdraws AI Report: 'Hallucinations' Question Reliability 2026-06-13
ToolSense: The Open-Source Framework for Evaluating LLM Tool Understanding
Frameworks ToolSense: The Open-Source Framework for Evaluating LLM Tool Understanding 2026-06-12
OpenAI Acquires Ona to Enhance Codex for Persistent Agent Work
LLM OpenAI Acquires Ona to Enhance Codex for Persistent Agent Work 2026-06-12
Deezer Introduces Tool to Identify AI-Generated Music on Streaming Platforms
Market Deezer Introduces Tool to Identify AI-Generated Music on Streaming Platforms 2026-06-11