Topic / Trend Rising

AI agents and process-oriented evaluation

New agent frameworks and benchmarks are shifting evaluation from final outputs to full trajectories, tool calls, errors and confidence. Projects span modular fact-checking, autonomous scientific agents, supply-chain optimization and self-repairing problem formulation.

Detected: 2026-09-15 · Updated: 2026-09-15

Related Coverage

2026-09-15 ArXiv cs.CL

PhysMent: the benchmark that turns LLMs into active experimenters

PhysMent evaluates LLMs on 105 classical mechanics scenes in the MuJoCo simulator: models must apply forces, query states, advance time, and modify geometry before answering. On quantitative multi-step tasks, most drop below 30 per cent: the bottlene...

#LLM On-Premise #Fine-Tuning #DevOps
2026-09-14 ArXiv cs.CL

R2VC: Modular Fact-Checking Splits Retrieval, Verification, and Confidence

R2VC introduces a modular fact-checking pipeline that separates Wikipedia retrieval, candidate generation, NLI verification, and confidence calibration. With an 8B backbone, it improves FEVER accuracy by 13.74% over baseline. Ablations show verifier-...

#Hardware #LLM On-Premise #Fine-Tuning
2026-09-12 ArXiv cs.AI

From Natural Language to QUBO: Iterative Self-Repair Makes the Difference

A multi-agent framework generates QUBO formulations from natural-language descriptions and uses test cases to correct them. On QUBOBench, 100 problems across 12 domains, it reaches 68% accuracy, beating a single-call baseline by 22%. The result shift...

#Hardware #LLM On-Premise #DevOps
2026-09-11 ArXiv cs.AI

OpenDiscoveryTrace: the missing process traces for AI scientist evaluation

A public dataset records 558 complete AI scientific agent trajectories with thoughts, tool calls, errors, and confidence. Traditional benchmarks look only at final outputs; here comparable frontier models reveal different error profiles. Claude Opus ...

#LLM On-Premise #Fine-Tuning #DevOps
← Back to All Topics