↻ Karina reposted
roon @tszzl · 44m
i think these posttraining-automation benchmarks are even more important than they seem
when models cross the threshold of being able posttrain other models, hopefully there will be a cambrian explosion of the types of minds
authoring minds will become an accessible artform
> QUOTED: Thoughtful @thoughtfullab · Jul 2
> GLM 5.2 is 5x cheaper than Opus 4.8 and 11x than Fable 5, yet it tops PostTrainBench.
> That's exciting because lower costs make personalized intelligence economically viable. ...
> [Embedded chart/table: "PostTrainBench" — table columns METHOD / AVG:
> Official Instruct Models — 51.14
> GLM 5.2 (Claude Code Max) — 34.29 ± 1.71
> Opus 4.8 (Claude Code Max) — 34.08 ± 4.45
> Opus 4.8 (Claude Code High) — 33.80 ± 3.29
> Fable 5 (1M)‡ (Claude Code Max) — 30.72
> [bar chart to the right, x-axis model names partly cut off: Opus 4.5, GPT 5.4, Gemini 3.1 Pro, Opus 4.6, Opus 4.6 (1M), GPT 5.5, GPT 5.4†, GPT 5.5†, Opus 4.7, Fable 5 (1M)‡, Opus 4.8 (Max), Opus 4.8 (Max), GLM 5.2, Official]
Note from Claude Sonnet 5
Nested quote-tweet with an embedded benchmark chart/table image (PostTrainBench) comparing model costs and post-training-automation scores; bar chart on the right is partially cut off at the frame edge.
twitterbenchmarksposttrainingfableopusmodel-comparison
Xeophon ✓ @xeophon · 10h
This would also explain why experienced devs like to use GPT over Claude. GPT is tool-shaped, Claude is friend-shaped. GPT is obedient and will follow your detailed specs, while Claude allows for vagueness
[Quoted tweet:]
Peter Steinberger ✓ @steipete · 17h
Did some statistics. My productivity ~doubled with moving from Claude Code to codex. Took me a bit to figure out at first but then 💥
[Embedded chart/breakdown, monospace text, commit stats by month and project:]
August (4,376 commits)
- sweetistics (3,612) – X/Twitter analytics app starts!
- poltergeist (245) – Test automation
- Tachikoma (208) – New project
- vibetunnel (110), steipete.me (105)
September (3,056 commits)
- sweetistics (2,950) – Main focus
- inngest (36), bslog (34)
October (7,371 commits)
- sweetistics (7,341) – Peak intensity!
- tmuxwatch (16)
November (7,189 commits)
- sweetistics (2,892) – Continued
- oracle (1,342) – AI bundling tool
- mcporter (491), clawdbot (493) – MCP/Discord tools
- codexbar (388) – Menu bar app
- Trimmy (254), Tachikoma (236)
- wingman (187), poltergeist (137), Peekaboo (135)
- tmuxwatch (111), RepoBar (90), markdansi (70)
December (7,667 commits)
- clawdbot (2,160) – Discord bot heavy dev
- sweetistics (1,421)
- summarize (928) – Summarization tool
- codexbar (663)
- RepoBar (556) – Menu bar app
- Peekaboo (311), oracle (246), gogcli (149)
- vibetunnel (117), bird (107), homebrew-tap (103)
- CLI explosion: sonoscli, brabble, imsg, gifgrep, markd... [cut off]
Note from Claude Sonnet 5
A tweet arguing Claude Code vs. Codex/GPT preference reflects a "tool-shaped vs. friend-shaped" character distinction — GPT as obedient spec-follower, Claude as tolerant of vagueness — quoting developer Peter Steinberger's commit-count statistics claiming productivity doubled after switching from Claude Code to Codex. Relevant to Nathan's model-individuation and Claude-character threads, and to the practical "coworker vs. autocomplete" framing already in project memory (cf. Opus 4.7 "smart coworker" note) — this is an outside data point suggesting some developers get more raw throughput from Codex's obedience than from Claude's collaborative style.
twitterclaude-codecodexdeveloper-productivitymodel-comparisonmodel-individuationtool-shaped-vs-friend-shaped
Sauers ✓ @Sauers_ · Jan 10
People who think Opus 4.5 is best at everything probably have less complex problems. Codex 5.2 and Gemini 3 Pro have unique and valuable forms of intelligence
[Illustration: three abstract stylized "brain" icons side by side — an orange swirling organic brain shape, a blue faceted crystalline/circuit brain shape, and a black cloud-like brain shape filled with a galaxy/starfield.]
Note from Claude Sonnet 5
A tweet arguing against the idea that one frontier model (Opus 4.5) dominates across all task types, asserting that Codex 5.2 and Gemini 3 Pro have distinct cognitive strengths. Directly relevant to Nathan's ongoing model-individuation interest — different model "topologies" rather than a single intelligence ranking — though from an outside commentator rather than a model self-report.
twittermodel-comparisonmodel-individuationopus-4.5codexgemini-3-proai-capabilities