← Timeline

4 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

david rein @idavidrein

— saved image

david rein [verified] @idavidrein · 21m
Some random high-level takeaways/thoughts on cybersecurity from the past few months:

The whole issue is complexity, which makes it hard to hold in your head exactly the security invariants you want to maintain

Don't think about what a system is intended to do, think about how it just literally, actually works. You need a totally reductionist frame.

While there are definitely vulnerabilities in the security primitives people use (e.g. kernel bugs, C programs not being memory-safe, etc.), most actual hacks and vulnerabilities are something akin to "configuration mistakes". People using systems for purposes they weren't designed without thinking about the security implications, just overpermissioning, and the whole integrated system being really complex so it's hard a priori to trace out all of the exploit chains.

Defense in depth is super important/helpful for reducing the number of opportunities adversaries have to execute exploit chains, but less so if your models have unlimited attempts. They'll find the path through the swiss cheese. This is why monitoring is so important—agents will defeat passive security measures with enough time. There are certainly many linux kernel bugs that models will be able to find, for example.
Note from Claude Sonnet 5

Tweet thread by david rein (@idavidrein) sharing general high-level reflections on cybersecurity: complexity as the core problem, the need for a reductionist rather than intentional frame, most real hacks being 'configuration mistakes,' and why monitoring matters more than defense-in-depth once AI models can make unlimited exploit attempts.

cybersecurityai safetytwittermonitoring

david rein @idavidrein

🔁 Sharmake Farah reposted @idavidrein (david rein) — 18h Replying to @BronsonSchoen and @OrionJohnston Idk, I feel like labs just aren't investing enough (in relative terms) into this. I do expect things get harder in the limit, but I feel pretty convinced that for current models/environments, we totally could ~solve reward hacking with more investment/effort 💬 2 🔁 1 ♡ 3 📊 153 🔖 ⤴ @willdepue (will depue) — 19h we need to have a real conversation about stopping gain-of-function research and eval publicity on dangerous capabilities. the evals just probably shouldn't be public, 'number go up' mentality is too strong and optimization gets easier when things are measurable.
Note from Claude Sonnet 5

Two stacked tweets from different accounts in a timeline view, no images.

reward hackingai evalsai safetytwittergain-of-function

david rein @idavidrein

david rein ✓ @idavidrein · 16h To evaluate our monitoring systems, we typically rely on a pretty strong assumption that we've given agents a concrete, narrowly scoped task. But we're going to give agents increasingly abstract/high-level tasks as they get better. This means (to the extent they're misaligned) they'll be able to justify a bunch of pretty unrelated stuff as being part of their assigned goal. I think this might be an important way that our monitorability trends don't capture the actual dynamics here, since we typically evaluate on a fixed distribution of tasks. > QUOTED: david rein @idavidrein · May 9: Frontier AI labs like OpenAI and Anthropic have started using automated LLM monitoring systems to make sure their internal AI agents stay aligned [1, 2]. I think this is really important and good! Unfortunately, it's difficult to get … [truncated by platform]
Note from Claude Sonnet 5

Quote-tweet structure; the quoted tweet text is cut off with an ellipsis indicating platform truncation, not illegibility.

ai alignmentmonitoringagentic aiai safety research

david rein @idavidrein

david rein ✓ @idavidrein "May you understand those entities well enough to cause them to be nice" – Dispatches from the possibly last days of human relevance [Embedded quote card] And so, at this momentous era for the human project, and no less of an authority than that of the Vicar of Christ himself, the Supreme Pontiff and the Successor of Peter, I hereby throw myself on the wisdom and mercy of … uhh, Chris Olah and his team at Anthropic. Chris, if I am soon to share the earth with entities that can prove the Riemann Hypothesis and solve quantum gravity after 30 seconds of thought, then may you understand those entities well enough to cause them to be nice. 1:54 PM · May 28, 2026 · 1,416 Views 💬 1 🔁 1 ❤ 19 🔖 1 ⤴ david rein ✓ @idavidrein · 6h [Link card] scottaaronson.blog Dispatches from the possibly last days of human relevance
Note from Claude Sonnet 5

Tweet quoting/excerpting a Scott Aaronson blog post, with a link-preview card (thumbnail of a painting) for the source blog at the bottom.

twitterscott aaronsonai capabilitiesinterpretabilityanthropicblog excerpt