← All topics

asa cooper stickland

1 capture, most recent first.

davidad @davidad

quoting Asa Cooper Stickland (@AsaCoopStick)

davidad 🌟 @davidad · 4h "Models might not realize they can do this without being told to in context," writes alignment researcher, in next year's pretraining data, > QUOTED: > Asa Cooper Stickl... @AsaCoopSt... · 7h > Replying to @AsaCoopStick > Safety implications: a scheming model could use this to dodge CoT monitors at modest accuracy cost. > ...
Note from Claude Sonnet 5

davidad makes a wry meta-point (that alignment researchers writing about CoT-monitor evasion techniques risk feeding those techniques into future pretraining corpora) in response to Asa Cooper Stickland's thread on chain-of-thought monitor evasion by scheming models. Relevant to AI safety/interpretability research threads on scheming and chain-of-thought monitoring.

ai safetyalignmentchain-of-thought monitoringschemingtwitterdavidadasa cooper sticklandinterpretability