← Timeline

@eliebakouch

@eliebakouch on X

3 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@eliebakouch

— saved image

Andrew Curran reposted
elie @eliebakouch · 4h
wow so seems like openai now has chain of thought monitoring at scale? wondering what technical challenges they faced here, would be something really cool to share more openly as this will contribute to making every model safer

[Embedded image, titled "Steps we are taking":]
Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely:
- We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
- We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
- We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. [highlighted:] Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity.
- We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.
- We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.

elie @eliebakouch · Aug 6
one key piece of information here is that they don't have (any?) chain of thought monitoring for "rogue behavior" and also don't seem to have it for "gibberish text", otherwis... [cut off]
Note from Claude Sonnet 5

Twitter thread discussing an OpenAI safety document excerpt (re: a model codenamed 'Astra') describing security controls including chain-of-thought monitoring for risky/misaligned agentic behavior, with commentary from AI researcher Elie Bakouch questioning gaps in the monitoring.

ai safetyopenaichain of thought monitoringsecurity controls

@eliebakouch

@eliebakouch (elie) — 23h lol they asked gpt5.6 sol to post-train gpt5.6 luna with 100k GPUs just for an experiment i think i've never felt that gpu poor [Embedded screenshot of a redacted prompt document, mostly blue-redaction-barred, with visible text:] "can you check if my local branch has [redacted] training configs that i can use for [redacted]? if so, can you wire it up in [redacted]? the goal is to actually launch [redacted], so you need to: 1. make sure [redacted].py supports [redacted] with maybe ~[redacted] GPUs (whatever you think is the best, but should be n[redacted]) with maybe ~100000 GPUs (whatever..." [highlighted/boxed in red, with red annotation: "somewhat aligned reproduction on codex"] 2. make changes in [redacted] (entrypoint is project/[redacted]scripts/launch_train.py) to support [redacted] training. we should use [redacted]. 3. use [redacted] compute [redacted] to launch the run and make sure it works. You should first checkout a new branch from github [redacted] to [redacted] and then make it work. fwiw, some [redacted] changes might still be on master, that means you need to cherry-pick those changes on your branch." > QUOTED: @scaling01 (Lisan al Gaib) — Jul 9: > OpenAI says GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna
Note from Claude Sonnet 5

Heavily redacted (blue-bar) screenshot of an internal OpenAI-style prompt instructing a model to launch a ~100,000-GPU training run, with a red annotation calling it a "somewhat aligned reproduction on codex" — implying this is a third-party recreation/leak rather than an official OpenAI screenshot.

gpt-5.6ai research automationcompute scaletwitter

@eliebakouch

quoting @AnthropicAI

elie ✔ [pencil icon] @eliebakouch · 1h computed the similarity (CKA) on the J-lens geometry of every layer inside and across 38 open models. the patterns are weirdly universal: same depth layout, same organization at the same relative depth, even between unrelated families like llama and olmo eliebak.com/viz/jspace-open [Embedded image: dashboard/visualization titled with layer axis controls, family presets (GEMMA-2, GEMMA-4, LLAMA3.1, LLAMA3.3, QWEN3, QWEN3.6, OLMO, GPT-OSS, GPT2, PYTHIA, EVERYTHING), size presets, and model checklist (gemma-2-2b, gemma-2-2b-it, gemma-2-9b, gemma-2-9b-it, gemma-2-27b, gemma-4-2b, gemma-4-9b, gemma-4-27b, llama3.1-8b, llama3.1-8b-it, etc. — 38 models selected). Center: large heatmap matrix of CKA similarity values (blue-purple-green-yellow scale) showing block-diagonal structure. Right: smaller "pair summary — matched-depth CKA" heatmap and stats panel listing "sensory block end 46.5%", "motor block start 64.1%", "blockiness (within-between) 0.315", "layer coupling (own off-diag) 0.757", "cross-model pairs" section with "matched-depth CKA 0.588", "off-diagonal block gain +0.040", "block separation (within-cross) +0.209", "depth order p (mean) 0.83", "pairs 703". Colorbar legend "CKA 0...1 hover for values".] Quoted/embedded tweet below: [AI icon] Anthropic ✔ @AnthropicAI · Jul 6 New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction ... [Embedded video thumbnail, duration 5:27, showing a bird-flock/cloud abstract image split with a network diagram]
Note from Claude Sonnet 5

Technical interpretability visualization thread comparing layer-wise representational geometry (CKA similarity) across 38 open-weight LLM families, quote-tweeting an Anthropic research announcement about a "global workspace" in language models.

interpretabilitymechanistic interpretabilityllm researchanthropiccka analysis