Topic / Trend Rising

Agentic AI, Evaluation and Security

Agentic AI is accelerating, but reliability and security remain central: LLM judges miss silent failures, MCP connectors expand the attack surface, open agentic models emerge, and high-stakes decision limits appear in oncology.

Detected: 2026-09-03 · Updated: 2026-09-03

Related Coverage

2026-09-03 ArXiv cs.CL

When silence is the right answer: training the evidence boundary

A training framework teaches grounded QA models to answer only when evidence becomes sufficient, locating the exact point where abstention gives way to a response. On HotpotQA, 2WikiMultiHopQA and MuSiQue, with Qwen2.5-3B-Instruct and LoRA, the metho...

#LLM On-Premise #Fine-Tuning #DevOps
2026-09-03 ArXiv cs.CL

PRO-Step: Step-Level Rewards for More Reliable Multi-Hop RAG

PRO-Step introduces step-level supervision for Retrieval-Augmented Generation. Instead of evaluating only the final answer, a generative process reward model checks logical validity and evidential grounding at every step of multi-hop reasoning. The m...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-02 ArXiv cs.CL

Outcome-Only LLM Judges Miss Silent Faults in Agent Trajectories

A study across 400 trajectories shows outcome-only LLM judges catch 84% of loud faults but only 45% of silent ones while flagging 33% of correct trajectories. A step-rubric judge reaches 77% silent recall with zero false alarms at 3x cost. No judge r...

#LLM On-Premise #DevOps
2026-09-01 ArXiv cs.LG

ERR+ rewards entropy drops for more efficient LLM reasoning

ERR+ introduces a two-phase RLVR framework that rewards token-level entropy drops and relative length efficiency. Correct traces show more frequent and larger entropy drops during thinking. Tests on five datasets report improved accuracy and concisen...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-01 ArXiv cs.AI

DS-Lighting: Making the Agent Harness Explicit for Data-Science Automation

DS-Lighting is an open-source toolkit that makes the agent harness explicit for data-science automation. It decomposes the harness into four reusable layers—data, workflow, execution, and evaluation—and represents agents as executable operator progra...

#Hardware #LLM On-Premise #DevOps
2026-08-28 LocalLLaMA

GLM-5.3: same base model, doubled cyber exploitation

Z.ai uses the same base model as GLM-5.2 for GLM-5.3, but all gains come from post-training. Coding improves by 50% on the internal Code Bench and reaches state of the art on Terminal Bench 3.0 and Agents' Last Exam. The real shift is the emergent cy...

#Hardware #LLM On-Premise #Fine-Tuning
2026-08-28 ArXiv cs.LG

NeuronFuzz: fuzzing that looks inside LLM safety neurons

A new white-box fuzzing framework replaces response-level feedback with a continuous score derived from safety neurons during prefill. Across 21 models it discovers jailbreaks in 76-100% of cases and transfers optimized templates to open-weight and p...

#LLM On-Premise #Fine-Tuning #DevOps
2026-08-27 LocalLLaMA

Apodex 1.1: open agentic models land in multiple quantized formats

The Apodex team's AMA on r/LocalLLaMA introduces Apodex 1.1, an open model family for agentic work, plus an open-source harness and two papers. Quantized variants in FP8, GPTQ-Int4, and NVFP4 point to local, self-hosted inference, but no benchmark fi...

#LLM On-Premise #DevOps
2026-08-27 LocalLLaMA

Apodex 1.1: open agentic models and quantization as a strategic signal

The Apodex team has presented Apodex 1.1, an open model family designed for complex work involving reasoning, search, code execution, and multi-agent coordination. Alongside the models, the team released FrontierAgent and two papers. The NVFP4, GPTQ-...

#Hardware #LLM On-Premise #DevOps
← Back to All Topics