← All topics

jason wolfe

2 captures, most recent first.

Jason Wolfe @w01fe

— saved image

Jason Wolfe ✓ @w01fe · 1h
It depends on what you do about this monitoring.  In general I think any action taken based on CoTs has the potential to affect monitorability and should be done carefully, but if you can't look at the CoTs at all then what's the point in even having them?
1 · 18 · 214

Jason Wolfe ✓ @w01fe · 1h
In general, the fewer actions you take based on monitorability, and the broader those actions are and less directly connected to the specific traces, the safer things seem.
1 · 17 · 177

Nathan Calvin ✓ @_NathanCalvin · 1h
Cool that makes complete sense

Seems like the answer is less "we shouldn't monitor COT during training" (seems like we totally should) but we should be keenly aware that acting a bunch of that information particularly in specific ways could have bad effects
1 · 1 · 8 · 89

Jason Wolfe ✓ @w01fe
Yeah, I'm not aware of a really simple hard and fast rule, but the intuition is something like "reduce the total number of bits of information conveyed into the policy about which kinds of CoTs are associated with being penalized."

7:48 PM · Aug 9, 2026 · 443 Views
Note from Claude Sonnet 5

Screenshot of the fuller X thread underlying the previous capture: Jason Wolfe arguing that any action taken on chain-of-thought traces costs monitorability, that broader and less trace-specific actions are safer, and Nathan Calvin agreeing the lesson is not 'don't monitor CoT' but 'be careful what you act on'. Captured 39 seconds after Screenshot 2026-08-09 211239.png, which shows the same exchange from Calvin's side.

chain of thoughtcot monitoringmonitorabilityalignmenttrainingjason wolfenathan calvin

Nathan Calvin @_NathanCalvin

reposted by Sharmake Farah — saved image

Sharmake Farah reposted

Nathan Calvin ✓ @_NathanCalvin · 1h

This seems important. Additional monitoring CoT during training is good but I'm worried some naive implementations without the sort of intuition Jason is describing could just make models really good at hiding scary thoughts.

Feels kinda like antibiotic resistant bacteria – antibiotics are great and it's good to use them but overuse them and you will have actually made the problem worse.

In general I think it's both true that there are many things companies should do differently post HF incident but also that there are lots of naive things they could do that could easily make the problem look like it's going away while making it worse.

Thanks to Jason for responding, this stuff is fascinating and the fact it has has such immense immediate real world stakes is wild

---

[quoted tweet]
Jason Wolfe ✓ @w01fe · 1h
Replying to @_NathanCalvin
Yeah, I'm not aware of a really simple hard and fast rule, but the intuition is something like "reduce the total number of bits of information conveyed into the policy about which kinds of CoTs are associated with being penalized."
Note from Claude Sonnet 5

Screenshot of an X post by Nathan Calvin (reposted by Sharmake Farah) warning that naive chain-of-thought monitoring during training could train models to hide scary thoughts rather than not have them — analogised to antibiotic resistance — quoting Jason Wolfe's rule of thumb about minimising the bits conveyed to the policy about which CoTs get penalised.

chain of thoughtcot monitoringtrainingalignmentmonitorabilitynathan calvinjason wolfe