← All topics

research automation

2 captures, most recent first.

Noam Brown @polynoamial

quoting @42_gravity (Chayanka_42)

Noam Brown ✓ @polynoamial · 58m This was one of the bigger open questions in quantum cryptography > QUOTED: Chayanka_42 ✓ @42_gravity · Jul 23 GPT-5.6 Sol Ultra helped solve a six-year-old open problem in quantum cryptography by generating the construction and main proof ideas through Codex. ... > [embedded white document excerpt, cut off]: AI usage. The human authors take full responsibility for the claims and pr[oofs] and have carefully refined and verified them. The construction and main idea[s were generate]d entirely by Codex using GPT 5.6 Sol Ultra, using harness ideas generated [by] [UCLA?] Moonshot Harness [ZHC+26] and [Ope26].
Note from Claude Sonnet 5

Screenshot showing a claim that GPT-5.6 Sol Ultra (via OpenAI Codex) generated the core construction/proof for a quantum cryptography research result; embedded paper "AI usage" disclosure statement is partially cut off at the right edge.

ai capabilitiesresearch automationquantum cryptographyopenaitwitter

Greg Brockman @gdb

quote-tweeting Jakub Pachocki (@merettm)

Greg Brockman @gdb · Feb 13 we are now benchmarking our models on novel frontier research, via firstproof.org. of 10 math research problems which research mathematicians have solved but never published the solutions to, in a week, our model discovered likely correct solutions to at least 6 of them. > QUOTED: Jakub Pachocki @merettm · Feb 13 > Very excited about the "First Proof" challenge. I believe novel frontier research is perhaps the most important way to evaluate capabilities of the next generation of AI models. > ... > Show more
Note from Claude Sonnet 5

OpenAI's Greg Brockman announcing "First Proof," a new benchmark testing AI models on unpublished, unsolved-in-literature research math problems — reporting their model found likely-correct solutions to 6 of 10 in a week. Relevant to Nathan's capability-progress tracking; a significant claimed jump in genuine novel-research capability rather than benchmark memorization.

openaibenchmarksmathematicsresearch automationcapability progressai timelinestwitter