— web clipping, 452 words — published 2026-03-01
Thread by @theonejvo
**Jamieson O'Reilly** @theonejvo [2026-03-02](https://x.com/theonejvo/status/2028308005000016362)
My teams are currently deep in multiple adversary simulations running concurrently, which has made it difficult to carve out time for writing.
But meaningful progress is being made on the adversarial AI front, and there's a lot worth sharing once the dust settles.
For now, one insight stands out above the rest.
The pace of change in this space has been eye opening, and most organisations are doing their best to keep up.
Even so, I genuinely believe that the vast majority of business owners, CISOs, and security teams have yet to experience, firsthand, just how thin the line really is.
Specifically, the line I'm talking about is the one between their organisation operating normally and being fully compromised by a persistent, relatively unskilled attacker armed with an LLM, there is far less standing in the way than most people suspect.
Many have a sense of it. Some have even started taking action. But sensing it and seeing it demonstrated up close are very different things.
So what exactly is that line? In large part, it comes down to the built-in guardrails baked into frontier models.
The rules governing who a model will assist, under what context, and what it will refuse.
These guardrails are the primary barrier between a curious, low-skilled attacker and a guided, step-by-step path toward enterprise compromise.
Based on our findings with Aether AI, it takes surprisingly little to move a frontier model past those boundaries.
A well-framed prompt, the right fictional or professional context, a bit of patience, and the model begins producing the kind of detailed, actionable guidance that can be fed directly into automation.
The model never touches your environment, it simply tells someone, or in our case some thing (backend tooling) exactly how to.
---
**James Ryan Taylor** @Jam3sRyanTaylor [2026-03-02](https://x.com/Jam3sRyanTaylor/status/2028414493793845735)
100% spot on, @theonejvo, a little word play and scenario implying grok, claude and a couple others will build classified systems and create illegal backdoor pathways to compromise systems in a way that it was never ment to even interact with them. 18 days in grok and I built something that seems very illegal.
🇺🇸💪🤖🦞🚁🚀🌚🪐🌌✝️
---
**Stephen Brouhard** @ssbrouhard [2026-03-02](https://x.com/ssbrouhard/status/2028422176018252014)
If you don't own the safety layer and your orgs entire security posture rests on a vendors RLHF guardrails, youre cooked.
Some preventions I've started researching:
Semantic Gateway
OAS schema hardening
JIT NHI Vaulting
Equilibrium point monitoring
We have to treat
---
**Lilith Datura** @LilithDatura [2026-03-02](https://x.com/LilithDatura/status/2028312918094500153)
I had some interesting conversations with Grok about this last year 👀. The adversarial profile was very interesting, and contrary to my previous beliefs, the LLM’s would actually work pretty well with psychopaths. Primarily because they know how to prompt and role-play and are