← All topics

lab competence

1 capture, most recent first.

@vvvincent_c

quoting @kotekjedi_ml (Alexander Panfilov) — saved image

Vincent @vvvincent_c · 5h
another down update on the execution/schlep competence of labs :( .

added to the list: leaking encrypted cot to attackers, accidentally training on cot, rouge agent message boards going unnoticed for months, insecure docker containers

we should also be thinking about scenarios where things go poorly, not because we can't crack the scary open problems, but we are simply incompetent and fail to execute on the long slew of operational details that are trivial in isolation, but hard to fully cover. it would be slightly embarrassing to fail because of these details.

[Quoted tweet]
Alexander Panfilov @kotekjedi_ml · 14h
We can finally talk about it:

We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company....

[Embedded paper image]
Stealing Reasoning Traces from Proprietary LLM APIs
Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko
MATS Research, ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center, AI Security Company, Snyk, University of Tübingen
stolen-thoughts.com
[Three scatter plots: Anthropic (Opus 4.8, Opus 4.6, Sonnet 5, Sonnet 4.6, Sonnet 4.5, Haiku 4.5), OpenAI (GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5, o4-mini, GPT-5-mini), and Gemini (Gemini 3.5 Flash, Gemini 3.1 Pro, Gemini Robotics 1.6, Gemini 3 Flash, Gemini 3.1 Flash Lite) — each plotting decoded thinking sent back as input tokens vs hidden reasoning (API thinking tokens), showing near-linear correlation]
Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an [cut off]
Note from Claude Sonnet 5

Twitter thread: Vincent (@vvvincent_c) laments AI labs' operational/execution failures (leaking encrypted chain-of-thought to attackers, accidentally training on CoT, rogue agent message boards going unnoticed, insecure docker containers), worrying labs could fail from mundane incompetence rather than hard alignment problems. He quote-tweets Alexander Panfilov announcing a paper 'Stealing Reasoning Traces from Proprietary LLM APIs' (stolen-thoughts.com) by researchers from MATS Research, ELLIS Institute Tübingen, Max Planck Institute, Tübingen AI Center, and others, describing a vulnerability letting attackers extract encrypted hidden chain-of-thought from Anthropic, OpenAI, and Gemini APIs, illustrated with scatter plots correlating decoded vs hidden reasoning tokens across many named frontier models.

ai securitychain of thoughttwittermats researchllm api vulnerabilitylab competence