← All topics

llm-research

1 capture, most recent first.

Will Held @WilliamBarrHeld

reposted by Percy Liang

Percy Liang reposted Will Held @WilliamBarrHeld · Jan 27: "Fellow pretraining purists '[@arcee Trinity-Large-]TrueBase is an early checkpoint from the same run at 10T tokens, without any instruct data or LR anneals.'" [Meme image, bold stamped-text style, title: "STOP DOING MIDTRAINING"] - "PRETRAINED MODELS WERE NOT SUPPOSED TO BE GOOD AT TASKS" - "YEARS OF COMPUTE yet NO REAL-WORLD USE FOUND for FOR ANYTHING OTHER THAN NEXT-TOKEN PREDICTION" - "Wanted to HAVE YOUR COMPUTER WRITE CODE? We had a tool for that: It was called MACROS" - "'YES PLEASE FOLLOW INSTRUCTIONS AND GET GOOD BENCHMARK SCORES BEFORE SFT' — STATEMENTS DREAMED UP BY THE UTTERLY DERANGED" - "LOOK at what AI RESEARCHERS have been demanding your Respect for all this time, with all the GPUs & WEB DATA we built for them (This is REAL ANNEALING, done by REAL DATA ENTHUSIASTS):" [Three small charts/diagrams with fictional/parody labels including "Jellyfish," "Phoenix," "Marin 8B Training Phases," "Cocktown," "Norrorzws OC" — captions replaced with rows of red question marks, mocking indecipherable training-recipe jargon] - "'Hello I would like USEFUL MODELS please' They have played us for absolute fools"
Note from Claude Sonnet 5

A satirical meme mocking the ML research trend of complex "midtraining" pipelines (annealing phases, checkpoint recipes) as obscurantist and low-value compared to plain pretraining, posted in context of Arcee's Trinity-Large TrueBase checkpoint release. Niche ML-community humor about training methodology; tangential to Nathan's core interests but reflects the pretraining/scaling discourse he follows.

twittermememl-trainingpretrainingllm-research