← All topics

safeguarded-ai

1 capture, most recent first.

davidad @davidad

quoting @ARIA_research

davidad 🌟✓ @davidad · 13h me@2024: Powerful AIs might all be misaligned; let's help humanity coordinate on formal verification and strict boxing me@2026: Too late! Powerful AIs are ~here, and some are open-weights. But some are aligned! Let's help *them* cooperate on formal verification and cybersecurity [Embedded chart: "Frontier AI's software R&D capabilities (Human-expert attention-span equivalent)" — a log-scale y-axis from 2s to 10h, x-axis quarterly from 2019 to 2028. Data points (x marks) trend upward from ~3s in 2019 to ~5h at a highlighted purple X around Q1 2026, with green/blue shaded uncertainty bands extending the trend to ~10-20h+ by 2027-2028. This appears to be a METR-style "time horizon" capability trend chart.] [Quoted tweet:] ARIA ✓ @ARIA_research · 15h In Safeguarded AI, we're funding teams to develop systems that harden our critical infrastructure from growing vulnerabilities. Programme Director @davidad warns that rapi... [cut off]
Note from Claude Sonnet 5

davidad (ARIA's Safeguarded AI programme director, known for formal-verification approaches to AI safety) reflects on how his own strategy shifted from 2024's "prevent misaligned AI" framing to 2026's "some AI is aligned, help it cooperate on verification/cybersecurity," alongside a METR-style capability trend chart showing frontier AI software R&D time-horizon capability crossing ~5 hours by early 2026. Directly relevant to Nathan's empirical singularity/AI-progress tracking notes (his memory already cites METR's 0.25-0.5 automation range and time-horizon doublings) — this chart is a current data point for that tracking thread, and davidad's strategic pivot is a notable AI-safety-field signal.

twitterdavidadariasafeguarded-aiformal-verificationai-capability-trendsmetrsingularity-trackingai-safety-strategy