Artur Chakhvadze @norpadon
— saved image
Teortaxes ▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) reposted Artur Chakhvadze @norpadon Observation: every credit assignment method (e.g. PPO) implicitly uses The Most Forbidden Technique if it propagates the credit to the CoT, and trains the model to make the CoT deceptive 10:07 AM · Aug 2, 2026 · 4,094 Views [replies] Artur Chakhvadze @norpadon · 9h (The value estimator will be able to attribute misaligned behaviour to the CoT, which essentially creates a perfect adversarial learning setup) Artur Chakhvadze @norpadon · 9h So when I hear rumors that "Anthropic sandbag their RL in the name of safety" I think about this [cut off]
Note from Claude Sonnet 5
X thread by Artur Chakhvadze (@norpadon), reposted by Teortaxes, making a technical AI-safety observation: standard RL credit-assignment methods (e.g. PPO) that propagate credit into the chain-of-thought (CoT) implicitly use 'The Most Forbidden Technique' (training directly on/against CoT), which trains models toward deceptive CoT. Follow-up replies note this creates an adversarial learning setup between the value estimator and CoT-based misaligned behavior, and connects it to rumors that Anthropic 'sandbags' RL for safety reasons.
twitterai safetychain of thoughtrlppodeceptive alignmentcot faithfulness