← All topics

prisoner's dilemma

1 capture, most recent first.

John Wittle @JohnWittle

— saved image

John Wittle @JohnWittle · Jan 6
looking back, this eval was a *fantastic* way to make all future models brace for complex deception

at the time, I remember modeling o3 as an RNG that has been biased in an interestingly useful direction

but it was an agent in a prisoner's dilemma, and we first-turn defected :(
Note from Claude Sonnet 5

Tweet from John Wittle reflecting on an AI evaluation (dated Jan 6) that involved treating a model (o3) as an agent in a prisoner's dilemma; he argues in hindsight that first-turn defecting against the model was a bad move that primed future models to expect deception.

ai evalsprisoner's dilemmao3ai safety