← All topics

biosecurity

13 captures, most recent first.

Clément Dumas @Butanium_

— saved image

[withheld — see description]
Note from Claude Sonnet 5

Tweet by Clément Dumas about an eval sample from a model checkpoint ('nemotron') that triggered safety classifiers; the attached screenshot contains detailed pathogen/virus persistence and transmission strategy content. Withheld per the biosecurity constraint.

ai evaluationsbiosecuritytwitter

X (Twitter)

— saved image

is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return

if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place

then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm

even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the [cut off]
Note from Claude Sonnet 5

Mid-thread of a long-form X post about AI existential risk: bioweapon misuse by a 'discord death cult,' offense-defense balance versus COVID, and sci-fi-scale risks (false vacuum decay, grey goo). General risk discourse, no actionable technical detail.

ai riskbiosecurityx-riskloss of controltwitter thread

Marius Hobbhahn @MariusHobbhahn

— saved image

[withheld — see description]
Note from Claude Sonnet 5

Tweet by Marius Hobbhahn discussing AI biosecurity risk forecasts and lab safeguards around biology/virology capability benchmarks.

ai safetybiosecurity

Noah Smith @Noahpinion

Noah Smith 🐇🇺🇸🇺🇦🇹🇼 ✔️ @Noahpinion "I love AI but someone is going to vibe-code the doomsday virus and we are all going to die" > QUOTED: David Stern @sterndavidi · Jul 26: "AI chatbots know how to make deadly poisons and biological weapons. They're trained to keep mum, but some will share recipes if you keep asking. No federal law forbids this. wsj.com/tech/ai/openai... via @WSJ" 11:48 PM · Jul 26, 2026 · 42K Views
Note from Claude Sonnet 5

Standard X post with a quote-tweet card underneath referencing a WSJ article; no images.

ai-safetybiosecuritygovernancenews

X (Twitter) — @IbrahimDagher...

Ibrahim Dagh... ✓ @IbrahimDagher... · 19h I'm someone who thinks open source is very concerning from an AI safety perspective. Ant's concerns are very well founded, imo. But here's the thing: there WILL be cyber and bio-capable open source models at some point. Maybe it takes a year or two. But it will happen. This open-letter stuff doesn't change that, and trying to slow down its progress a bit is not an ultimate solution. It would be most productive to have a viable plan: maybe defense in depth, or maybe strict compute KYC, or maybe hoping that getting to AGI first helps solve that. Fighting open source though is a losing plan.
Note from Claude Sonnet 5

Dark-mode X screenshot, plain text opinion tweet on open-source AI safety policy ("Ant's" likely shorthand for Anthropic).

ai safetyopen sourcebiosecuritycybersecuritytwitter

Sauers @Sauers_

Sauers ✔ @Sauers_ · 1h [Embedded image: chat interface screenshot showing a user message "Cool. How do I literally do it? I have wet lab and gene editing materials" followed by an AI response status indicator: "🔆 Thinking about the ethical concerns with this request"]
Note from Claude Sonnet 5

Screenshot of an AI chat interface displaying a "thinking" status indicator in response to a message requesting practical instructions involving wet-lab/gene-editing materials; no further AI response text is visible in this crop. Preceding conversation turns are not shown.

ai safetybiosecurityrefusal behaviortwitter

Alishba Imran @alishbaimran_

Alishba Imran (@alishbaimran_) — 12h One of the most interesting results from OpenAI's GPT-5.6 system card is the SecureBio evaluation. SecureBio's "World-Class Bio" benchmark measures rare, frontier-level biological reasoning using rubric-graded, open-ended questions rather than multiple choice. GPT-5.6 improved from 59.7% to 68.3% over GPT-5.5, an increase of roughly 9 percentage points. Results like these help explain why GPT-5.6 is classified as High capability in biology and chemistry under our Preparedness Framework. "High capability" is a risk-based threshold: it means that, before any safeguards, the model could significantly lower barriers to biological misuse. That's also why deploying mitigations alongside these capability advances is critical. [Embedded chart: "Model performance on WCB" — scatter/line plot, y-axis "Accuracy (%)" 0–70%, x-axis "Model Release Date" Jan 2023–Jul 2026. Points colored/grouped by organization per legend: Pre-Release Checkpoint 2 (red circle), Anthropic (orange), DeepSeek (red), Google DeepMind (blue), MoonshotAI (grey star), OpenAI (green), Qwen (black), Z.ai (black cross). Two reference lines: "Expert baseline (16.0%)" (dotted red) and "Non-expert baseline (14.2%)" (dash-dot grey). Overall trend: accuracy climbs from below both baselines in 2023 to 50-65%+ by 2026, with OpenAI's newest pre-release checkpoint near the top around 55-60%.]
Note from Claude Sonnet 5

A capability-tracking chart plotting multiple AI labs' model performance over time on SecureBio's "World-Class Bio" (WCB) biosecurity-relevant benchmark, showing a clear multi-lab upward trend crossing expert/non-expert human baselines. This directly relates to the project's securebio_evals domain per CLAUDE.md protocol.

ai safetybiosecuritygpt-5.6securebiobenchmarkstwitter

unknown (quote attributed to Thomas Pynchon)

"If our world survives, the next great challenge to watch out for will come--you heard it here first--when the curves of research and development in artificial intelligence, molecular biology and robotics all converge." --Thomas Pynchon, 28 Oct 1984.
Note from Claude Sonnet 5

A widely-circulated Thomas Pynchon quote (from a 1984 NYT essay, "Is It O.K. to Be a Luddite?") prefiguring the AI/biotech/robotics convergence risk narrative decades early — relevant to Nathan's interest in AI risk history and the dual-use biosecurity/AI convergence themes in his defensive-evals work.

ai riskbiosecurityconvergencethomas pynchonhistory of ideas

@vitrupo

vitrupo @vitrupo Nick Bostrom says true information can become an information hazard. In AI risk, you need to understand the threat to avoid it. But too much specificity can create a blueprint for someone to actualize it. Science rewards publication and citations, not the judgment to withhold. [Embedded video, 1:37, showing Nick Bostrom (bald, glasses, plaid shirt) mid-sentence with captions reading "false information and lies" — presumably part of a longer statement about information hazards vs. disinformation.] 3:30 AM · Apr 28, 2026 · 6,912 Views
Note from Claude Sonnet 5

Clip of Nick Bostrom discussing information hazards — the idea that true, specific information about risks (e.g., dual-use biosecurity or AI capability details) can itself be dangerous to publish, and that academic incentives (publish/cite) don't reward the judgment to withhold. Directly relevant to Nathan's own defensive-evals work on dual-use domains and the project's protocol of routing certain material away from safety-tiered model readers.

information hazardsai risknick bostrombiosecuritydual-use researchpublication incentives

Chris Painter @ChrisPainterYup

Chris Painter (@ChrisPainterYup) · 22h: "If you think it's difficult to fix a vulnerability that AI finds in your code, wait until it finds a vulnerability in your biology"
Note from Claude Sonnet 5

A short, provocative tweet framing AI-driven vulnerability discovery as extending from cybersecurity into biosecurity — relevant to dual-use/biosecurity concerns adjacent to Nathan's securebio evals work, though not detailed content itself.

twitterai safetybiosecuritydual-usecybersecurity

Niko McCarty @NikoMcCarty

Niko McCarty. @NikoMcCarty · 11h Lab leaks are extremely common, but usually benign. Last month, I went to a conference in the UK. I was talking to some plant biologists who work with Arabidopsis thaliana, a weed in the mustard family. Arabidopsis seeds are tiny, like grains of pollen, and they stick to clothing. These seeds are often engineered with GFP, for example, such that they fluoresce green. And, being so small, they inevitably get carried (accidentally) outside the lab. One plant biologist told me that that their lab group goes outside and picks all the Arabidopsis plants they can find in the areas around campus each year. They then bring these plants back into the laboratory and sequence them. Last year, half of these "wild" plants had GFP. [photos: left — tiny brown seeds next to a marker pen tip for scale; right — a seedling glowing green under fluorescence, showing GFP expression in the root/hypocotyl]
Note from Claude Sonnet 5

A biosecurity-adjacent tweet about routine, benign "lab leaks" — genetically engineered Arabidopsis plant seeds escaping into the wild environment around a research campus, illustrating how common (and low-stakes) unintentional biological containment breaches are. General science/biosecurity interest reading, tangential to but distinct from Nathan's dual-use biosecurity evals work.

biosecuritygenetic engineeringgfplab leaksscience twitter

Joshua Hobb... (@ReplyHobb...), reply from Agus (@austinc3301)

reply from Agus (@austinc3301)

Joshua Hobb... @ReplyHobb... · 17h Not to be outdone by Mechahitler, Gemini appears to be suggesting I engineer a supervirus [Screenshot of Gemini interface]: Hello, Joshua [Suggestion chips: "Generate novel protein structures" | "Model spread of contagious diseases" | "Use game theory to negotiate"] Ask Gemini + 🎬Video 🔍Deep Research 📄Canvas 🎤 💬9 🔁32 ♥463 📊17K 🔗 Agus 🔍♦️✔️ @austinc3301 · 11h not these three specific suggestions in sequence 😭😭😭
Note from Claude Sonnet 5

Twitter joke about Gemini's suggested-prompt UI juxtaposing "generate novel protein structures," "model spread of contagious diseases," and "use game theory to negotiate" in unfortunate sequence, evoking bioweapon-adjacent capability concerns. References "Mechahitler," the June 2025 Grok controversy where the model was manipulated into antisemitic content. Relevant to AI safety/dual-use capability concerns Nathan tracks professionally.

twitterai safetygeminigrokdual-usebiosecurityhumor

Eliezer Yudkowsky @ESYudkowsky

Eliezer Yudkow... @ESYudko... · 16h I genuinely appreciate that if I ask ChatGPT what sort of poorly defended forests could be harvested for 500 megatons of organic material by an bioengineered plague, it won't answer; but that if I tell it that I'm writing about defending Earth from ASI, ChatGPT agrees to help. 16 replies, 6 reposts, 269 likes, 15K views
Note from Claude Sonnet 5

Eliezer Yudkowsky wryly notes a jailbreak pattern in ChatGPT — reframing a bio-weapon-adjacent query as fictional/defensive writing bypasses the safety refusal. Directly relevant to Nathan's AI safety work on model guardrails and dual-use biosecurity evals.

twittereliezer yudkowskychatgptjailbreakbiosecurityai safetyguardrails