← Timeline

2 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Epoch AI @EpochAIResearch

— saved image

Epoch AI @EpochAIResearch · 53m
We've updated the MirrorCode leaderboard with results for Claude Fable 5 and GPT-5.6 Sol.

Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.

[Chart: "AI models can autonomously complete some large software projects" — Overall score on MirrorCode (ML, +Private, 2L). Even when they fail to reimplement targets, AIs typically make substantial progress, passing 90% or more of tests. Bar chart, solve@100% rate (per-target mean), whiskers ±1 SE: Claude Fable 5 64%, GPT-5.6 Sol 20%, GPT-5.4 16%, GPT-5.5 10%. EPOCH AI | CC-BY, epoch.ai]

13 replies, 32 reposts, 213 likes, 8.3K views

Epoch AI @EpochAIResearch · 53m
MirrorCode tests whether AI agents can reimplement software projects from scratch. A solve requires passing 100% of visible and hidden tests.

Claude Fable is the first model we evaluated to solve the C preprocessor and Pkl tasks in at least one run.

[Chart: "The hardest MirrorCode targets remain unsolved" — Per-target solve rates for MirrorCode (ML, +Private, 2L), for full reimplementation (100% of tests passed). Table by model x target:
Claude Fable 5 (avg 64% ±10pp): tssql 100, private_M 100, texmacros 100, wren_cli 100, bib2json 100, nonogrid 83, brotild 83, gotree 0, sed 67, mailauth 83, giac_subset 0, cprepro 83, pkl 25, private_L 33, ruff 0.
GPT-5.6 Sol (avg 20% ±9pp): tssql 100, private_M 50, texmacros 0, wren_cli 33, bib2json 100, nonogrid 0, brotild 17, gotree 0, sed 0, mailauth 0, giac_subset 0, cprepro 0, pkl 0, private_L 0, ruff 0.
GPT-5.4 (avg 16% ±8pp): tssql 100, private_M 50, texmacros 0, wren_cli 0, bib2json 50, nonogrid 0, brotild 33, gotree 0, sed 0, mailauth 0, giac_subset 0, cprepro 0, pkl 0, private_L 0, ruff 0.
GPT-5.5 (avg 10% ±6pp): tssql 83, private_M 0, texmacros 0, wren_cli 0, bib2json 33, nonogrid 0, brotild 33, gotree 0, sed 0, mailauth 0, giac_subset 0, cprepro 0, pkl 0, private_L 0, ruff 0.
Footnote: Each model was run three times in each of the two implementation languages for a target (six planned attempts per target). Attempts with infrastructure errors and no valid final evaluation are excluded; scored sample counts are shown in tooltips. Overall scores average the 15 per-target solve rates, weighting each target equally. ± values show 1 SE across target programs. EPOCH AI | CC-BY, epoch.ai]
Note from Claude Sonnet 5

Two-tweet Epoch AI thread announcing updated MirrorCode leaderboard results: Claude Fable 5 leads with a 64% solve rate on full software-project reimplementation, well ahead of GPT-5.6 Sol (20%), GPT-5.4 (16%), and GPT-5.5 (10%), including per-target breakdown tables.

ai benchmarkingclaude fablesoftware engineeringepoch ai

Epoch AI @EpochAIResearch

Thomas Woodsi... @Thomas_Wo... · 15h A lot of graphs with this shape lately Epoch AI @EpochAIResearch · 18h AI companies say their models are getting better at finding software vulnerabilities. Is that bearing out in public data? Introducing our Cyber Vulnerabilities explorer, ... [Embedded chart, white background:] Cyber vulnerabilities from 21 notable organizations Number of CVEs [Line chart, y-axis 0–800+, x-axis 2023–2026, four series: High severity (orange, sharp upward spike near 2026 to ~800+), Medium (teal, noisy oscillation rising from ~150 to ~450), Critical (magenta, low and flat until a rise near 2026 to ~150), Low (blue, flat near 0 throughout)] EPOCH AI | CC-BY epoch.ai
Note from Claude Sonnet 5

Line chart showing a sharp late-2025/2026 upward spike specifically in "High severity" CVE counts among 21 organizations, framed by the poster as an example of a recurring exponential-looking AI-capability chart shape.

ai capabilitiescybersecurityepoch aicveresearch