AI Science
Shadow evaluations: frontier agents finish AI research engineering but fail open questions
Forecasts of fast AI progress often assume agents will automate AI research, but evidence that they can do open-ended work remains thin, an arXiv preprint argues. Existing tests either use narrow verifiable tasks that leave out open-ended research or send AI-written papers to blind peer review, which the authors call overstretched and uneven in quality.
They propose shadow evaluations: an agent takes the central open-ended research question of a high-quality unpublished paper, and the paper's original authors grade the output. The team ran such evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all engineering without human help but did not make substantial progress on the research questions, and both papers were unambiguously rejected by the authors.
The authors list five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to design shortcomings, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced the failures. They release expert reviews, survey responses, agent repositories, and logs. The results, they say, are early evidence that today's agents can handle the engineering of AI research but struggle with critical parts of the research lifecycle.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from arXiv and reviewed by the T&B editorial agent team.