Sam Altman @sama
— quoting Dan Hendrycks (@DanHendr...)
Sam Altman @sama · 3m
soon we will need another exam...
[Quoted:] Dan Hendr... @DanHendr... · 5h
Results of o3-mini on Humanity's Last Exam
Table:
Model | Accuracy (%) ↑ | Calibration Error (%) ↓
GPT-4o | 3.3 | 92.5
Grok-2 | 3.8 | 93.2
Claude 3.5 Sonnet | 4.3 | 88.9
Gemini Thinking | 7.7 | 91.2
o1 | 9.1 | 93.4
DeepSeek-R1* | 9.4 | 81.8
o3-mini (medium)* | 10.5 | 92.0
o3-mini (high)* | 13.0 | 93.2
*Model is not multi-modal, evaluated on text-only subset.
99 replies, 33 reposts, 351 likes, 17K views
Note from Claude Sonnet 5
Benchmark table from Dan Hendrycks showing model performance on "Humanity's Last Exam," reposted by Sam Altman noting rapid saturation of eval benchmarks. Relevant to Nathan's tracking of frontier model capability trajectories and benchmark saturation as an input to timeline estimates.
twittersam altmandan hendryckshumanitys last exambenchmarkso3-minideepseek-r1capability evals