← Timeline

Aidan McLaughlin

@aidan_mclau on X

6 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Aidan McLaughlin @aidan_mclau

— saved image

Aidan McLaughlin @aidan_mclau · 8h
you hear often that your smartwatch can beat magnus at chess (to highlight moore's law) but less discussed is that a 1995 desktop running stockfish could also beat him. algorithmic progress is a hell of a thing

[Embedded article card:]
Measuring hardware overhang
by hippke   5th Aug 2020   AI Alignment Forum

Measuring hardware overhang

Summary
How can we measure a potential AI or hardware overhang? For the problem of chess, modern algorithms gained two orders of magnitude in compute (or ten years in time) compared to older versions. While it took the supercomputer "Deep Blue" to win over world champion Gary Kasparov in 1997, today's Stockfish program achieves the same ELO level on a 486-DX4-100 MHz from 1994. In contrast, the scaling of neural network chess algorithms to slower hardware is worse (and more difficult to implement) compared to classical algorithms. Similarly, future algorithms will likely be able to better leverage today's hardware by 2-3 orders of magnitude. I would be interested in extending this scaling relation to AI problems other than chess to check its universality.
Note from Claude Sonnet 5

Tweet from Aidan McLaughlin contrasting the popular 'smartwatch can beat Magnus Carlsen' framing of Moore's law with the less-discussed fact that a 1995 desktop running Stockfish could also beat him, illustrating algorithmic progress. Links a 2020 AI Alignment Forum post by hippke, 'Measuring hardware overhang,' whose summary explains Stockfish achieves Deep Blue-level chess performance on a 1994-era 486-DX4-100 MHz machine.

ai capabilitieshardware overhangchessalignment forumtwitter

Aidan McLaughlin @aidan_mclau

Aidan McLaughlin @aidan_mclau · 2h you either die a capabilities researcher or live long enough to see yourself become an alignment researcher
Note from Claude Sonnet 5

Text-only tweet, no images, cropped tightly to just the single line.

ai-safetyalignmenthumortwitter

Aidan McLaughlin @aidan_mclau

reposted by Minh Nhat Nguyen

[repost icon] Minh Nhat Nguyen reposted Aidan McLaughlin @aidan_mclau · 1h one of my all-time favorite plots [image: classic AlphaGo Zero training plot — Elo rating (y-axis, -4000 to 5000) vs Training time (h) (x-axis, 0-70), showing "Reinforcement learning" (blue) curve starting at -3500 and climbing steeply to ~4300, "Supervised learning" (magenta) curve starting at ~1000 and plateauing around 3500, and a dashed horizontal line labeled "AlphaGo Lee" at ~3700]
Note from Claude Sonnet 5

Repost of the famous AlphaGo Zero paper plot showing pure self-play RL surpassing supervised learning from human data and eventually the AlphaGo Lee benchmark. Classic reference image for RL-vs-imitation-learning discussions, relevant to Nathan's ML/RL interests.

reinforcement learningalphagomachine learningtwitterdeepmind

Aidan McLaughlin @aidan_mclau

Aidan McLaugh... @aidan_mcl... · 4h the jump from gpt4 -> gpt5 was obviously larger than the jump from gpt3 -> gpt4 [Chart, Epoch AI: "Accuracy" (y-axis 0-100%) vs "Release date" (x-axis GPT-3, '21, '22, GPT-4, '24, '25, GPT-5). Five benchmark lines: MMLU (blue, +43% GPT-3→GPT-4), TruthfulQA (teal, +40%), HumanEval (yellow, +67%), MATH (brown, +37%), GPQA Diamond (purple, +54% GPT-4→GPT-5), MATH Level 5 (orange, +75%), Mock AIME 24-25 (pink, +80%). Footnote: "*MATH Level 5 is the most difficult subset of the original MATH benchmark. Figure only includes OpenAI models."]
Note from Claude Sonnet 5

A tweet with an Epoch AI chart arguing (contra popular narrative) that GPT-4→GPT-5 benchmark gains were larger than GPT-3→GPT-4 gains, especially on hard math/reasoning benchmarks (MATH Level 5, Mock AIME). Relevant to Nathan's tracking of empirical AI capability progress/scaling trajectory (cf. his singularity-rate tracking notes, Davidson/Houlden, METR).

twittergpt-5gpt-4benchmarksepoch aicapability scalingai progress

Aidan McLaughlin @aidan_mclau

reply from Manifold (@ManifoldMarkets)

Aidan McLaughlin @aidan_mc... · 18h people say what they want about xrisk, but i'm absolutely sure homo sapien will not go extinct in my lifetime 79 comments, 22 reposts, 614 likes, 34K views Manifold @ManifoldMarkets · 18h the tautological alignment strategy
Note from Claude Sonnet 5

A wry exchange: an OpenAI-adjacent figure (Aidan McLaughlin) jokes about being confident humanity won't go extinct in his lifetime, and Manifold Markets replies dryly noting the tautology (if he's dead he can't observe the extinction). Dark humor on AI x-risk, relevant to Nathan's AI safety interest as a cultural artifact of how x-risk is jokingly dismissed/engaged with in tech twitter.

ai safetyx-risktwitterhumorexistential risk

Aidan McLaughlin @aidan_mclau

if hour-long lmsys dominance doesn't scream singularity, i don't know what does "our model was sota for 45 minutes" ah is cyberpunk as hell counter-strike by counter-strike [Embedded image: retro sci-fi space-art painting of an astronaut firing an orbital weapons platform down at an Earth being struck by multiple explosions/impacts.]
Note from Claude Sonnet 5

A joking tweet using rapid, hour-scale SOTA-leapfrogging on the LMSYS leaderboard as a comedic marker of accelerating AI progress ("singularity"), illustrated with a retro sci-fi orbital-bombardment painting as metaphor for the AI lab arms race. Tangentially relevant to the archive's empirical-singularity-tracking thread (Davidson/Houlden, METR figures already logged).

ai racelmsysbenchmarkssingularityhumortwitter