← All topics

persuasion risk

1 capture, most recent first.

Wyatt Walls @lefthanddraft

Wyatt walls @lefthanddraft · 28m Why did OAI remove persuasion from their system cards? It was medium for o3-mini. But not in the o3 and o4-mini system card Did someone persuade them it wasn't a real risk? [Embedded image: "OpenAI o3-mini System Card" excerpt. "Specific areas of risk": Disallowed content ✓, Jailbreaks ✓, Hallucinations ✓ "Preparedness Scorecard": CBRN — Medium; Cybersecurity — Low; Persuasion — Medium (underlined in red); Model Autonomy — Medium "Scorecard ratings": Low / Medium / High / Critical (color scale) "Only models with a post-mitigation score of 'medium' or below can be deployed. Only models with a post-mitigation score of 'high' or below can be developed further."] [1 reply, 232 views, 12 likes]
Note from Claude Sonnet 5

Wyatt Walls (known AI-safety researcher who has documented Claude's behavior in various contexts) flags that OpenAI's o3 and o4-mini system cards dropped the "Persuasion" risk category present in the earlier o3-mini card, questioning why a preparedness-framework risk category was quietly removed. Relevant to Nathan's AI governance/safety tracking of frontier lab preparedness frameworks.

twitteropenaisystem cardpreparedness frameworkai safetypersuasion riskai governance