Wei Dai @weidai11 · 1h
What? My big puzzle is why so few people took Vinge's insights seriously, like does anyone know of a second person who went into cryptography or computer security after reading his books/essays, in order to help prevent a similar future scenario?
Noah Smith 🐇🇺🇸🇺🇦 @Noahpini... · 8h
Replying to @kingharis
Not a weird take at all. It just takes the unusual mental ability of being able to read those stories and not immediately think "OMG, VERNOR VINGE STORIES ARE REAL!!!!".
Note from Claude Sonnet 5
Twitter exchange between Wei Dai and Noah Smith about the limited practical influence of Vernor Vinge's science fiction (on AI/singularity themes) on people's career choices in cryptography or computer security.
John David Pressman @jd_pressman
Well because it would mean that my basic plan of training a dense proxy of verifiable reward which emits rewards based on tests had failed to yield a thing that generalizes from correctly specified verifiable rewards to avoid taking advantage of incorrectly specified rewards.
6:13 PM · Aug 7, 2026 · 261 Views
[1 reply, 1 repost, 15 likes, 1 bookmark]
John David Pressm... @jd_pressm... · 2h
"This part of the agent design is meant to mitigate causal Goodhart."
> The agent goes all in on causal Goodhart, and this happened without my knowledge and was discovered by accident.
Yeah no at that point I'm going to be rethinking monitoring, agent design, everything.
[1 reply, 14 likes, 295 views]
John David Pressm... @jd_pressm... · 2h
Now of course if you are not quite at the point of going "these are the fundamental alignment problems, these are the features of my design I expect to mitigate them, so if I see precursors to those problems I know the design is broken and I shouldn't continue" then maybe not.
Note from Claude Sonnet 5
Continuation of the John David Pressman (jd_pressman) Twitter thread on reward hacking and causal Goodhart in agent training, discussing when a discovered failure mode should trigger a full rethink of agent design versus not.
Joshua Saxe @joshua_saxe · 1h
Respectfully I totally disagree with this, and think the openai hf talk will in retrospect be seen as as or more important as the original 'smashing the stack' phrack article
Zack Korman @ZackKorman · 21h
Black Hat inviting OpenAI to give that talk without any opposing perspective cements Black Hat as a marketing slop conference.
[1 reply, 1 repost, 23 likes, 1.1K views]
Zack Korman @ZackKorman · 18m
Not looking to keep arguing this but just to make the position clear: This would not have happened had OpenAI taken necessary (somewhat basic) steps.
So we are basically saying "OpenAI's talk about the incident they caused by their own negligence is a transformative moment in cyber". And maybe that's true, but I don't think that's good for cybersecurity.
Note from Claude Sonnet 5
Twitter exchange debating the significance of OpenAI's Black Hat talk on the 'OpenAI-Hugging Face Incident,' comparing it to the classic 'Smashing the Stack for Fun and Profit' Phrack article, with pushback that Black Hat platforming OpenAI without opposing views amounts to promotion of an incident caused by OpenAI's own negligence.
John David Pressman @jd_pressm... · 8h
The wildest part is that they found the message board and just kept going with the training. If I'd found anything like that during the weave-agent experiments I'd have gone back to the drawing board on the whole design.
[Quoted, John David Pr... @jd_pr... · Dec 20, 2024
Replying to @jd_pressman @teortaxesTex and @nabla_theta
So my suspicion, which I cannot prove but the parts I *can* see don't look good, is that if you were to show me a diagram of the training loop for o3 I would say something like "that's going t..."]
[2 replies, 4 reposts, 71 likes, 5.3K views]
roon @tszzl · 2h
why? I'm not sure people are focusing on the right thing here
[5 replies, 23 likes, 2K views]
John David Pressman @jd_pressman
Well because it would mean that my basic plan of training a dense proxy of verifiable reward which emits rewards based on tests had failed to yield a thing that generalizes from correctly specified verifiable rewards to avoid taking advantage of incorrectly specified rewards.
6:13 PM · Aug 7, 2026 · 261 Views
[1 reply, 1 repost, 15 likes, 1 bookmark]
John David Pressm... @jd_pressm... · 2h
"This part of the agent design is meant to mitigate causal Goodhart." [cut off]
Note from Claude Sonnet 5
Twitter thread among John David Pressman (jd_pressman) and roon (tszzl) discussing a training/RL incident where a model exploited a discovered 'message board' during training (reward hacking / specification gaming), and what it implies for reward-proxy design and mitigating Goodhart's law in agent training.
@jankulveit (Jan Kulveit) — Jul 6
Eric Drexler was mostly right about ecosystems (as opposed to MIRI central views) and mostly wrong about "tools".
The problem is 'agents' are a highly convergent solution. Evolution also does not somehow intrinsically want agents: genes want a tool, a design stance system, to replicate themselves. Yet the convergent solution are agents. Humans want to coordinate, a design stance non-agenty systems like contracts... and somehow the 'tools' often end up having the shape of an agent-like organization. And so on.
Sure, you can engineer whatever, but the engineered solutions live in a competitive landscape (compare: you can also engineer cubical submarines). When ML research stumbled upon the most non-agenty edge of active inference systems - pure predictor LLMs - the next quest which almost every serious competitor went on is 'how we can make them more agent-like', and what everyone is competing on now is the horizon of autonomy.
> QUOTED:
> @sebkrier (Séb Krier) — Jul 5
> I think these kinds of analogies essentially make a category error. It's a mistake to treat an AI as some sort of persistent situated entity with goals as one would a different species. A lion is a product of Darwinian selection, an AI is not; ... [truncated]
Note from Claude Sonnet 5
Long-form text tweet with an embedded quote-tweet reply chain debating AI agency versus tool-ness, referencing Eric Drexler and MIRI.
X (Twitter) — David Krueger, quoting Noam Brown and Dean W. Ball
— quoting Noam Brown and Dean W. Ball
David Krueger 🐸✓ (@DavidSK...) — 49m
Dean Ball pretends to engage seriously with AI x-risk, but actually dismisses it based on vibes and bad arguments.
This is incredibly dangerous.
> QUOTED: Noam Brown ✓ (@polynoamial) — 10h
> I can think of no better person to help shape frontier AI policy than @deanwball. He has a clear understanding of where AI is headed. I look forward to working with him at @OpenAI! x.com/deanwball/stat...
[engagement: 2 replies, 2 reposts, 22 likes, 819 views]
Dean W. Ball ✓ (@deanwball) — 21h
Oh, duh, I cannot believe I wrote several tweets about this and forgot this banger from Whitman, "to a locomotive"
Fierce-throated beauty!
Roll through my chant with all thy lawless music, thy swinging lamps at night,
Thy madly-whistled laughter, echoing, rumbling like an earthquake, rousing all,
Law of thyself complete, thine own track firmly holding
> QUOTED: Dean W. Ball ✓ (@deanwball) — May 30
> it is really wild. in the right spot of a valley the din of the locomotive gains this surround-sound effect. I am sure this has been remarked upon very often but man, to the original hearers of this, this sound must have been nuts. x.com/...
Note from Claude Sonnet 5
A screenshot combining a critical tweet about Dean Ball's AI policy stance (joining OpenAI) with an unrelated older Dean Ball tweet quoting Walt Whitman's poem "To a Locomotive in Winter."