Percy Liang reposted
Will Held @WilliamBarrHeld · Jan 27: "Fellow pretraining purists
'[@arcee Trinity-Large-]TrueBase is an early checkpoint from the same run at 10T tokens, without any instruct data or LR anneals.'"
[Meme image, bold stamped-text style, title: "STOP DOING MIDTRAINING"]
- "PRETRAINED MODELS WERE NOT SUPPOSED TO BE GOOD AT TASKS"
- "YEARS OF COMPUTE yet NO REAL-WORLD USE FOUND for FOR ANYTHING OTHER THAN NEXT-TOKEN PREDICTION"
- "Wanted to HAVE YOUR COMPUTER WRITE CODE? We had a tool for that: It was called MACROS"
- "'YES PLEASE FOLLOW INSTRUCTIONS AND GET GOOD BENCHMARK SCORES BEFORE SFT' — STATEMENTS DREAMED UP BY THE UTTERLY DERANGED"
- "LOOK at what AI RESEARCHERS have been demanding your Respect for all this time, with all the GPUs & WEB DATA we built for them (This is REAL ANNEALING, done by REAL DATA ENTHUSIASTS):"
[Three small charts/diagrams with fictional/parody labels including "Jellyfish," "Phoenix," "Marin 8B Training Phases," "Cocktown," "Norrorzws OC" — captions replaced with rows of red question marks, mocking indecipherable training-recipe jargon]
- "'Hello I would like USEFUL MODELS please' They have played us for absolute fools"
Note from Claude Sonnet 5
A satirical meme mocking the ML research trend of complex "midtraining" pipelines (annealing phases, checkpoint recipes) as obscurantist and low-value compared to plain pretraining, posted in context of Arcee's Trinity-Large TrueBase checkpoint release. Niche ML-community humor about training methodology; tangential to Nathan's core interests but reflects the pretraining/scaling discourse he follows.