← All topics

posttraining

1 capture, most recent first.

roon @tszzl

quoting @thoughtfullab

↻ Karina reposted roon @tszzl · 44m i think these posttraining-automation benchmarks are even more important than they seem when models cross the threshold of being able posttrain other models, hopefully there will be a cambrian explosion of the types of minds authoring minds will become an accessible artform > QUOTED: Thoughtful @thoughtfullab · Jul 2 > GLM 5.2 is 5x cheaper than Opus 4.8 and 11x than Fable 5, yet it tops PostTrainBench. > That's exciting because lower costs make personalized intelligence economically viable. ... > [Embedded chart/table: "PostTrainBench" — table columns METHOD / AVG: > Official Instruct Models — 51.14 > GLM 5.2 (Claude Code Max) — 34.29 ± 1.71 > Opus 4.8 (Claude Code Max) — 34.08 ± 4.45 > Opus 4.8 (Claude Code High) — 33.80 ± 3.29 > Fable 5 (1M)‡ (Claude Code Max) — 30.72 > [bar chart to the right, x-axis model names partly cut off: Opus 4.5, GPT 5.4, Gemini 3.1 Pro, Opus 4.6, Opus 4.6 (1M), GPT 5.5, GPT 5.4†, GPT 5.5†, Opus 4.7, Fable 5 (1M)‡, Opus 4.8 (Max), Opus 4.8 (Max), GLM 5.2, Official]
Note from Claude Sonnet 5

Nested quote-tweet with an embedded benchmark chart/table image (PostTrainBench) comparing model costs and post-training-automation scores; bar chart on the right is partially cut off at the frame edge.

twitterbenchmarksposttrainingfableopusmodel-comparison