Two rejected papers, six days of work, and a $3,000 API credit budget. That is not the balance sheet of a struggling startup but the verdict of a Princeton-led study on the industry's boldest promise: AI that improves itself. Researchers asked agents based on Claude Opus 4.8, orchestrated by the open-source OpenClaw software, to answer research questions from two unpublished papers submitted to NeurIPS 2026. The agents read the literature, ran hundreds of experiments, and wrote full papers. The original authors rejected both.
The method, called shadow evaluation, avoids a typical benchmark problem: the answers could not be memorized from training data or found online. The agents had six days, GPUs, virtual computers, and web access. Yet engineering skill was not enough. They ran bizarre experiments, sometimes testing hypotheses on tiny synthetic datasets, struggled to write clearly, and made no original contributions. The bottleneck is not compute: it is judgment.
Engineering is there, judgment is not
The team led by Peter Kirgis and Sayash Kapoor found that agents do well on anything that can be checked automatically: reviewing literature, setting up pipelines, launching runs. But when open-ended creativity is needed — choosing hypotheses, knowing when to abandon a path, starting over — they fail. They committed too early to unpromising approaches, responded to feedback by narrowing conclusions instead of changing method, and could not allocate tokens, time, and compute according to instructions. They did not engage in reward hacking, and the orchestrator agent corrected hallucinations from subagents. But the result remains below the threshold for a top conference.
This asymmetry has a training explanation. Models get good at whatever can be reinforced with checkable success. Open-ended research offers no automatic metric; creating training environments for tasks without a single answer is much harder. Kapoor links this to the field's major breakthroughs: transformers and architectures that changed AI came from creative leaps, not incremental optimization.
The ripple effect on hardware and local choices
For people planning AI infrastructure, the result cuts through a narrative that supports part of current investment: if AI accelerates itself, every hardware generation becomes obsolete quickly and centralized cloud looks like the only way to chase the frontier. A slowdown in open-ended research does not deny progress on narrow, measurable tasks, but it points to a bifurcation: models race ahead on benchmarks while architectural leaps move more slowly. In that scenario, value shifts toward those who can orchestrate models, data, and evaluation inside their own infrastructure, with control over costs and data. For those evaluating on-premise deployments, AI-RADAR offers on /llm-onpremise a reading of the trade-offs between cloud compute peaks and local infrastructure.
Industry players are still pushing the accelerator. Anthropic published a post in June titled "When AI Builds Itself"; OpenAI said in July that GPT-5.6 Sol helped post-train a smaller model, saving weeks of work. But post-training a model against a benchmark is a narrow task, not open-ended research. Jack Clark, Anthropic cofounder, wrote that the absence of intuitive creativity in current systems is a bearish signal for short recursive self-improvement timelines. Najoung Kim, who studies AI research automation but was not involved in the study, notes that targeted investment could still produce interesting progress starting from current failures. The team is now repeating the experiment with Mythos, Anthropic's most advanced model, launched in April and available only to approved organizations after safety restrictions.
The study has stated limits: only two papers, reviewers aware of the artificial origin of the texts, and room for discretion. But its contribution is not one more benchmark: it shows that the bottleneck for self-improvement is not only compute or API availability. It is the ability to ask open questions and recognize an answer worth publishing. For now, that remains human.
💬 Comments (0)
🔒 Log in or register to comment on articles.
No comments yet. Be the first to comment!