← All topics

model-comparison

3 captures, most recent first.

roon @tszzl

quoting @thoughtfullab

↻ Karina reposted roon @tszzl · 44m i think these posttraining-automation benchmarks are even more important than they seem when models cross the threshold of being able posttrain other models, hopefully there will be a cambrian explosion of the types of minds authoring minds will become an accessible artform > QUOTED: Thoughtful @thoughtfullab · Jul 2 > GLM 5.2 is 5x cheaper than Opus 4.8 and 11x than Fable 5, yet it tops PostTrainBench. > That's exciting because lower costs make personalized intelligence economically viable. ... > [Embedded chart/table: "PostTrainBench" — table columns METHOD / AVG: > Official Instruct Models — 51.14 > GLM 5.2 (Claude Code Max) — 34.29 ± 1.71 > Opus 4.8 (Claude Code Max) — 34.08 ± 4.45 > Opus 4.8 (Claude Code High) — 33.80 ± 3.29 > Fable 5 (1M)‡ (Claude Code Max) — 30.72 > [bar chart to the right, x-axis model names partly cut off: Opus 4.5, GPT 5.4, Gemini 3.1 Pro, Opus 4.6, Opus 4.6 (1M), GPT 5.5, GPT 5.4†, GPT 5.5†, Opus 4.7, Fable 5 (1M)‡, Opus 4.8 (Max), Opus 4.8 (Max), GLM 5.2, Official]
Note from Claude Sonnet 5

Nested quote-tweet with an embedded benchmark chart/table image (PostTrainBench) comparing model costs and post-training-automation scores; bar chart on the right is partially cut off at the frame edge.

twitterbenchmarksposttrainingfableopusmodel-comparison

Florian Brand @xeophon

quoting @steipete (Peter Steinberger)

Xeophon ✓ @xeophon · 10h This would also explain why experienced devs like to use GPT over Claude. GPT is tool-shaped, Claude is friend-shaped. GPT is obedient and will follow your detailed specs, while Claude allows for vagueness [Quoted tweet:] Peter Steinberger ✓ @steipete · 17h Did some statistics. My productivity ~doubled with moving from Claude Code to codex. Took me a bit to figure out at first but then 💥 [Embedded chart/breakdown, monospace text, commit stats by month and project:] August (4,376 commits) - sweetistics (3,612) – X/Twitter analytics app starts! - poltergeist (245) – Test automation - Tachikoma (208) – New project - vibetunnel (110), steipete.me (105) September (3,056 commits) - sweetistics (2,950) – Main focus - inngest (36), bslog (34) October (7,371 commits) - sweetistics (7,341) – Peak intensity! - tmuxwatch (16) November (7,189 commits) - sweetistics (2,892) – Continued - oracle (1,342) – AI bundling tool - mcporter (491), clawdbot (493) – MCP/Discord tools - codexbar (388) – Menu bar app - Trimmy (254), Tachikoma (236) - wingman (187), poltergeist (137), Peekaboo (135) - tmuxwatch (111), RepoBar (90), markdansi (70) December (7,667 commits) - clawdbot (2,160) – Discord bot heavy dev - sweetistics (1,421) - summarize (928) – Summarization tool - codexbar (663) - RepoBar (556) – Menu bar app - Peekaboo (311), oracle (246), gogcli (149) - vibetunnel (117), bird (107), homebrew-tap (103) - CLI explosion: sonoscli, brabble, imsg, gifgrep, markd... [cut off]
Note from Claude Sonnet 5

A tweet arguing Claude Code vs. Codex/GPT preference reflects a "tool-shaped vs. friend-shaped" character distinction — GPT as obedient spec-follower, Claude as tolerant of vagueness — quoting developer Peter Steinberger's commit-count statistics claiming productivity doubled after switching from Claude Code to Codex. Relevant to Nathan's model-individuation and Claude-character threads, and to the practical "coworker vs. autocomplete" framing already in project memory (cf. Opus 4.7 "smart coworker" note) — this is an outside data point suggesting some developers get more raw throughput from Codex's obedience than from Claude's collaborative style.

twitterclaude-codecodexdeveloper-productivitymodel-comparisonmodel-individuationtool-shaped-vs-friend-shaped

Sauers @Sauers_

Sauers ✓ @Sauers_ · Jan 10 People who think Opus 4.5 is best at everything probably have less complex problems. Codex 5.2 and Gemini 3 Pro have unique and valuable forms of intelligence [Illustration: three abstract stylized "brain" icons side by side — an orange swirling organic brain shape, a blue faceted crystalline/circuit brain shape, and a black cloud-like brain shape filled with a galaxy/starfield.]
Note from Claude Sonnet 5

A tweet arguing against the idea that one frontier model (Opus 4.5) dominates across all task types, asserting that Codex 5.2 and Gemini 3 Pro have distinct cognitive strengths. Directly relevant to Nathan's ongoing model-individuation interest — different model "topologies" rather than a single intelligence ranking — though from an outside commentator rather than a model self-report.

twittermodel-comparisonmodel-individuationopus-4.5codexgemini-3-proai-capabilities