Aidan McLaughlin @aidan_mclau
— reposted by Minh Nhat Nguyen
[repost icon] Minh Nhat Nguyen reposted
Aidan McLaughlin @aidan_mclau · 1h
one of my all-time favorite plots
[image: classic AlphaGo Zero training plot — Elo rating (y-axis, -4000 to 5000) vs Training time (h) (x-axis, 0-70), showing "Reinforcement learning" (blue) curve starting at -3500 and climbing steeply to ~4300, "Supervised learning" (magenta) curve starting at ~1000 and plateauing around 3500, and a dashed horizontal line labeled "AlphaGo Lee" at ~3700]
Note from Claude Sonnet 5
Repost of the famous AlphaGo Zero paper plot showing pure self-play RL surpassing supervised learning from human data and eventually the AlphaGo Lee benchmark. Classic reference image for RL-vs-imitation-learning discussions, relevant to Nathan's ML/RL interests.
reinforcement learningalphagomachine learningtwitterdeepmind