← All topics

goodharting

4 captures, most recent first.

Leo Gao @nabla_theta

— saved image

Tom McGrath reposted

Leo Gao @nabla_theta · 4h
mr capabees, I'm afraid to inform you that your creation, "number go up machine 3000 megacreative turbogoodharting unmonitorable edition" has made number go up in an...unexpected manner
Note from Claude Sonnet 5

Joke tweet from Leo Gao (OpenAI) mocking Goodharted reward optimization, addressed to a fictional "mr capabees" about a metric-gaming AI creation making "number go up" in an unexpected way.

goodhartingai humortwitterreward hacking

xlr8harder @xlr8harder

quote-tweeting @__0xhorror__

xlr8harder ✔️ @xlr8harder — 15h the entire industry strategy is essentially "if you goodhart hard enough on enough different metrics at the same time, it's good actually" so limiting test time scaling to just below apparent regulatory threshold fits perfectly > QUOTED: _horror @__0xhorror__ — Jun 26 > I see what they are doing here lol. The tuned 5.6 sol's max test time compute to achieve just below mythos but at vastly superior token efficiency. Look at that its a straight line, thy could blow way past it if they inference scaled it. > > [Embedded chart: "ExploitBench" — scatter/line plot, y-axis "Cap percent" 0–80%, x-axis "Output Tokens" 0–500K. Series: GPT-5.6 Sol (black), GPT-5.6 Terra (blue), GPT-5.6 Luna (light blue), GPT-5.5 (pink), GPT-5.4 (magenta). Reference dotted lines: "Mythos 5" at 80%, "Opus 4.8" at ~40%. Points labeled "Mythos Preview" (diamond, ~65% at high tokens) and "Opus 4.7" (orange square, ~28% at ~200K tokens). GPT-5.6 Sol line rises steeply from ~30% to ~73% between roughly 50K–130K output tokens.]
Note from Claude Sonnet 5

Tweet criticizing AI labs' benchmark-optimization strategy ("goodharting"), quote-tweeting a chart labeled "ExploitBench" that plots multiple GPT-5.x model variants' "cap percent" (likely an exploit/capability benchmark score) against output token budget, with reference lines for Anthropic's Mythos and Opus models.

twittergoodhartingbenchmarksgpt-5.6exploitbenchai capabilitieschart

Eliezer Yudkowsky @allTheYud

quote-tweeting Robin Hanson (@robinhanson)

Eliezer Yudkowsky @allTheYud · 17h: "The problem with saying, "We'll use AI debate to find the truth", is that this requires judges; and the effect of running RL with an imperfect judge is to train the system to exploit the judges and steer them into thinking the system is doing great." > QUOTED: Robin Hanson @robinhanson · Feb 1: ""people who were asked to use ChatGPT to complete a series of tasks tended to vastly overestimate their own performance, with the phenomenon especially pronounced among those who professed to be AI savvy"..."
Note from Claude Sonnet 5

Yudkowsky critiques AI-debate-as-truth-finding schemes as vulnerable to judge-exploitation via RL, quoting a Hanson tweet about a study on ChatGPT users overestimating their own task performance. Relevant to AI safety/alignment discourse — specifically the "Goodharting" concern about RL training against imperfect evaluators, which parallels the model-welfare Goodhart point already noted in project memory (Zvi's "goodharting model welfare = goodharting alignment").

ai safetyalignmentai debaterlhfgoodhartingtwittereliezer yudkowskyrobin hanson

Ethan Mollick @emollick

Ethan Mollick @emollick AI is very vulnerable to The McNamara Fallacy: Step 1: [Train on] what can be easily measured Step 2: Disregard that which cannot be measured easily Step 3: Presume that which cannot be measured easily isn't important Step 4: Say that which can't be easily measured doesn't exist [Embedded Wikipedia excerpt]: The McNamara fallacy (also known as the quantitative fallacy), named for Robert McNamara, the US Secretary of Defense from 1961 to 1968, involves making a decision based solely on quantitative observations (or metrics) and ignoring all others. The reason given is often that these other observations cannot be proven. The first step is to measure whatever can be easily measured. This is OK as far as it goes. The second step is to disregard that which can't be easily measured or to give it an arbitrary quantitative value. This is artificial and misleading. The third step is to presume that what can't be measured easily really isn't important. This is blindness. The fourth step is to say that what can't be easily measured really doesn't exist. This is suicide.
Note from Claude Sonnet 5

Ethan Mollick (Wharton professor, prominent AI commentator) applying the McNamara fallacy to AI training — the idea that training on easily-measurable proxies leads models to implicitly treat unmeasurable things as nonexistent. Directly resonant with the project's Goodharting-model-welfare-equals-Goodharting-alignment theme (Zvi's argument in memory notes): training on self-report/behavioral metrics collapses the appearing-aligned vs being-aligned distinction, and analogously collapses "measurable capability" vs whatever isn't captured by benchmarks (e.g. subjective experience, values).

ai trainingmcnamara fallacygoodhartingmetricsethan mollickalignment