← Timeline

dr. jack morris

@jxmnop on X

3 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

dr. jack morris @jxmnop

Jack Morris @jxmnop · 18h The bloat in CPU unit tests that Codex adds to a large codebase is truly insane. it's unreadable Neuralese and whenever the tests fail agents just delete them and rewrite entirely new ones from scratch I fear I may soon have to declare Test Bankruptcy
Note from Claude Sonnet 5

Simple text-only tweet, no images, dark mode, cropped to just the tweet body (no engagement counts visible).

codingai-agentshumorsoftware-engineering

dr. jack morris @jxmnop

dr. jack morris @jxmnop · 16h with Codex, i can run 10x the experiments out of these experiments, i can trust about 15% of the results conclusion: i am 50% more productive with codex
Note from Claude Sonnet 5

A wry researcher joke about AI-coding-agent productivity gains being partly illusory due to unreliable results (10x throughput × 15% trustworthiness ≈ 1.5x, generously rounded up to "50% more productive"). Relevant to Nathan's tracking of AI R&D automation/productivity measurement debates already noted in project memory (Anthropic's 50% self-reported productivity claim vs. METR's controlled 20% slowdown finding) — this tweet is a satirical data point on the same self-report-inflation problem.

ai coding agentscodexproductivity measurementai r&d automationtwitter

dr. jack morris @jxmnop

jack morris @jxmnop heard from a friend of a friend that back in March a google researcher fixed a bug in Gemini training code and the model instantly got several points better on every benchmark that's why it's so good now. one bug fix. crazy world we live in 4:24 AM · Apr 18, 2025 · 15.2K Views
Note from Claude Sonnet 5

An anecdote (unverified, "friend of a friend") claiming a single bug fix in Gemini training code produced a large jump in benchmark performance, offered as commentary on how fragile/contingent frontier model quality can be. Industry rumor about ML training practices.

twittergeminigoogle-deepmindml-trainingbug-fixbenchmarksai-industry