← All topics

ai-research-automation

1 capture, most recent first.

Kevin A. Bryan @Afinetheorem

quote-tweeting @sayashk (Sayash Kapoor); linked paper via CRUX

Kevin A. Bryan (@Afinetheorem) — 11:05 AM · Jul 30, 2026 · 20.5K Views Great work here from an all-star team on where we stand on RSI via AI research. My mental model is "what year will AI independently come up with an idea as valuable as Chinchilla Law or MoE". 2026: not yet. But again, folks I ask this Q give me 2027 as the modal answer... > QUOTED: @sayashk (Sayash Kapoor) — Jul 30 > Can AI agents conduct open-ended AI research? > Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, ... [truncated by platform] > [Embedded paper card image, teal background:] > "Can AI agents conduct open-ended AI research? Early evidence from two case studies" > Authors: Peter Kirgis*†, Sayash Kapoor*†, Andrew Schwartz, Stephan Rabanser†, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan > Affiliations: 1 Princeton University, 2 Cornflower Labs, 3 UK AI Security Institute, 4 University of Toronto, 5 Independent, 6 UC Berkeley, 7 Georgetown University (CSET), 8 Johns Hopkins University, 9 Golden Gate Institute for AI, 10 AI Digest, 11 Stanford University > * Equal contribution † CRUX Core Team > Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle. > Date: July 30, 2026 > Reproduction materials: https://cruxevals.com > [ALT badge, CRUX logo]
Note from Claude Sonnet 5

Twitter discussion of a new "shadow evaluations" paper (CRUX/Princeton/UKAISI et al.) testing whether frontier AI agents can conduct open-ended AI research; both test papers were rejected by their original authors despite agents finishing all engineering work. Directly relevant to the project's recursive self-improvement / singularity-timeline tracking thread.

twitterai-research-automationrsisingularity-timelinearxiv-paper