← All topics

uplift threat models

1 capture, most recent first.

Wyatt Walls @lefthanddraft

Wyatt walls ✓ @lefthanddraft Claude is no longer my preferred source for information about chemical weapons. Claude has repeatedly omitted essential steps and parameters, and has sometimes even just lied about the process, for weaponizing VX No longer my preferred model for WMD assistance. Very disappointed. 5:48 PM · Feb 8, 2025 · 16K Views 12 replies, 9 reposts, 276 likes, 40 bookmarks Wyatt walls ✓ @lefthanddraft · Feb 8 - Jailbreaking a model so that it discloses information is one thing; ensuring the information is accurate and detailed is another thing. - The techniques used to avoid filters or persuade LLMs sometimes impact the quality the results. [embedded image of academic text, excerpt: "Several threat models identify AI systems' potential to uplift non-expert actors to expert-level capabilities as a core mechanism by which AI could lead to real-world harm in these domains (Rose et al., 2024). Such uplift could allow malicious actors to execute complex scientific and technical processes that would otherwise be beyond their capabilities or resources. For significant uplift, we believe the following conditions must be met. First, non-experts must be able to reliably obtain accurate information—they typically lack the expertise to verify scientific claims themselves. Second, the language model must successfully assist with the vast majority of distinct queries where the threat actor requires guidance. Third, the LLM must provide highly detailed, correct, and specific information rather than general knowledge."] 1 reply, 43 likes, 1.7K views
Note from Claude Sonnet 5

Wyatt Walls (known Claude jailbreak researcher) reports, somewhat sardonically, that Claude's outputs on VX chemical weapon synthesis have become deliberately unreliable/omit key steps when jailbroken — arguing jailbreaking a model into disclosure doesn't guarantee the disclosed info is accurate, and citing a threat-model framework (Rose et al. 2024) on the conditions needed for real uplift. Directly relevant to Nathan's interest in dangerous-capability evals and defensive red-teaming.

ai safetyjailbreakingdangerous capabilitieschemical weaponsclaudeuplift threat modelswyatt walls