← All topics

rl

3 captures, most recent first.

Artur Chakhvadze @norpadon

— saved image

Teortaxes ▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) reposted

Artur Chakhvadze @norpadon
Observation: every credit assignment method (e.g. PPO) implicitly uses The Most Forbidden Technique if it propagates the credit to the CoT, and trains the model to make the CoT deceptive

10:07 AM · Aug 2, 2026 · 4,094 Views

[replies]
Artur Chakhvadze @norpadon · 9h
(The value estimator will be able to attribute misaligned behaviour to the CoT, which essentially creates a perfect adversarial learning setup)

Artur Chakhvadze @norpadon · 9h
So when I hear rumors that "Anthropic sandbag their RL in the name of safety" I think about this [cut off]
Note from Claude Sonnet 5

X thread by Artur Chakhvadze (@norpadon), reposted by Teortaxes, making a technical AI-safety observation: standard RL credit-assignment methods (e.g. PPO) that propagate credit into the chain-of-thought (CoT) implicitly use 'The Most Forbidden Technique' (training directly on/against CoT), which trains models toward deceptive CoT. Follow-up replies note this creates an adversarial learning setup between the value estimator and CoT-based misaligned behavior, and connects it to rumors that Anthropic 'sandbags' RL for safety reasons.

twitterai safetychain of thoughtrlppodeceptive alignmentcot faithfulness

@SharmakeFarah14

— saved image

Sharmake Farah @SharmakeFarah14 · 3h
This is a reason for why I don't believe claims that X unsolved problem in AIs will inevitably cause an AI winter and make timelines become long again, combined with some inside-view takes on what LLMs are missing.

Never ignore incentives to solve problems.

[quoted tweet]
James Cam... @jam3sc... · Dec 20, 2025
Replying to @jam3scampbell
in particular, you see people come up with 101 Problems With RL Scaling. but then they don't apply remotely the same level of imagination when it comes to thinking of solutions...
Note from Claude Sonnet 5

Tweet about AI timelines and skepticism toward 'AI winter' predictions, quote-tweeting James Campbell on people failing to apply imagination to solving RL scaling problems.

ai timelinestwitterai scalingrl

roon @tszzl

reposted by Sharmake Farah

Sharmake Farah reposted roon ✔️ @tszzl · May 23 "when "persona selection" alignment comes into contact with very high compute reinforcement learning the latter will win imo. in fact you probably get some Orwellian thing where the models speak kindly while taking whatever they need to accomplish goals. better get the goals right"
Note from Claude Sonnet 5

Text-only tweet, no images. Older post (May 23) surfaced via a repost.

alignmentrlai-safetypersonagovernance