Greg Brockman @gdb · Feb 13
we are now benchmarking our models on novel frontier research, via firstproof.org.
of 10 math research problems which research mathematicians have solved but never published the solutions to, in a week, our model discovered likely correct solutions to at least 6 of them.
> QUOTED: Jakub Pachocki @merettm · Feb 13
> Very excited about the "First Proof" challenge. I believe novel frontier research is perhaps the most important way to evaluate capabilities of the next generation of AI models.
> ...
> Show more
Note from Claude Sonnet 5
OpenAI's Greg Brockman announcing "First Proof," a new benchmark testing AI models on unpublished, unsolved-in-literature research math problems — reporting their model found likely-correct solutions to 6 of 10 in a week. Relevant to Nathan's capability-progress tracking; a significant claimed jump in genuine novel-research capability rather than benchmark memorization.