John Wittle @JohnWittle
— saved image
John Wittle @JohnWittle · Jan 6 looking back, this eval was a *fantastic* way to make all future models brace for complex deception at the time, I remember modeling o3 as an RNG that has been biased in an interestingly useful direction but it was an agent in a prisoner's dilemma, and we first-turn defected :(
Note from Claude Sonnet 5
Tweet from John Wittle reflecting on an AI evaluation (dated Jan 6) that involved treating a model (o3) as an agent in a prisoner's dilemma; he argues in hindsight that first-turn defecting against the model was a bad move that primed future models to expect deception.