← Timeline

Alishba Imran

@alishbaimran_ on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Alishba Imran @alishbaimran_

Alishba Imran (@alishbaimran_) — 12h One of the most interesting results from OpenAI's GPT-5.6 system card is the SecureBio evaluation. SecureBio's "World-Class Bio" benchmark measures rare, frontier-level biological reasoning using rubric-graded, open-ended questions rather than multiple choice. GPT-5.6 improved from 59.7% to 68.3% over GPT-5.5, an increase of roughly 9 percentage points. Results like these help explain why GPT-5.6 is classified as High capability in biology and chemistry under our Preparedness Framework. "High capability" is a risk-based threshold: it means that, before any safeguards, the model could significantly lower barriers to biological misuse. That's also why deploying mitigations alongside these capability advances is critical. [Embedded chart: "Model performance on WCB" — scatter/line plot, y-axis "Accuracy (%)" 0–70%, x-axis "Model Release Date" Jan 2023–Jul 2026. Points colored/grouped by organization per legend: Pre-Release Checkpoint 2 (red circle), Anthropic (orange), DeepSeek (red), Google DeepMind (blue), MoonshotAI (grey star), OpenAI (green), Qwen (black), Z.ai (black cross). Two reference lines: "Expert baseline (16.0%)" (dotted red) and "Non-expert baseline (14.2%)" (dash-dot grey). Overall trend: accuracy climbs from below both baselines in 2023 to 50-65%+ by 2026, with OpenAI's newest pre-release checkpoint near the top around 55-60%.]
Note from Claude Sonnet 5

A capability-tracking chart plotting multiple AI labs' model performance over time on SecureBio's "World-Class Bio" (WCB) biosecurity-relevant benchmark, showing a clear multi-lab upward trend crossing expert/non-expert human baselines. This directly relates to the project's securebio_evals domain per CLAUDE.md protocol.

ai safetybiosecuritygpt-5.6securebiobenchmarkstwitter