← All topics

misalignment-score

1 capture, most recent first.

Miles Brundage @Miles_Brundage

quoting Jan Leike (@janleike); replies from binal (@binalkp91) and J (@tofdaj)

Miles Brundage ✓ @Miles_Brundage Isn't there a company you forgot to mention 🧐 > QUOTED: Jan Leike ✓ @janleike · 3h > Interesting trend: models have been getting a lot more aligned over the course of 2025. > The fraction of misaligned behavior found by automated auditing has been going down not just at ... > [Embedded chart: scatter plot, x-axis dates 2023-01 to 2025-10+, y-axis "Misalignment Score" 0.00–0.20+. Points colored by company: Anthropic (tan) — Opus 3, Opus 4, Opus 4.5, Haiku 4.5; OpenAI (red) — GPT-4, o1, o3, GPT-5, GPT-5.2; Google (blue) — Gemini 2.0 Flash, Gemini 2.5 Pro, Gemini 3 Pro; xAI (black) — Grok 3, Grok 4.1. Trend lines per company: OpenAI (red, dashed) declining steadily from ~0.18 to ~0.10; Anthropic (tan, dashed) declining from ~0.16 to ~0.11; Google (blue, dashed) declining sharply from ~0.20 (Gemini 2.0 Flash) to ~0.11 (Gemini 3 Pro); xAI (black, dashed) the only company trending UP, from ~0.17 (Grok 3) to ~0.20 (Grok 4.1).] 12:08 PM · Jan 20, 2026 · 4,716 Views 💬4 🔁1 ❤️34 🔖4 [share] binal ✓ @binalkp91 · 1h "always up and to the right" 💬 🔁 ❤️5 📊278 🔖 [share] J @tofdaj · 1m xAI drives AI safety by showing what happens when you don't care about AI safety (Can't remember whose quote)
Note from Claude Sonnet 5

Jan Leike (former OpenAI/Anthropic alignment lead) shares a chart showing that "misalignment score" (fraction of misaligned behavior found by automated auditing) has trended down for OpenAI, Anthropic, and Google models through 2025, but xAI's models (Grok 3 → Grok 4.1) are the sole outlier trending upward; Miles Brundage jabs that Leike "forgot to mention" xAI in his original framing. Directly and highly relevant to Nathan's AI safety/alignment tracking — a concrete empirical trend on cross-lab misalignment measurement, with xAI singled out as diverging.

alignmentjan-leikemisalignment-scoreautomated-auditingxaigrokopenaianthropicgoogle-deepmindmiles-brundagetwitterai-safetycross-lab-comparison