About

Nathan Helm-Burger

Research Manager at MATS (Machine Alignment, Transparency, and Security), where I help early-career researchers develop technical AI safety research agendas.

Background

I started in neuroscience — PhD program at Georgetown studying human intelligence enhancement. But I'd been following AI research for years, and eventually concluded that timelines were too short for human intelligence enhancement to improve our chances at alignment. So I switched to working on AI directly, spending several years in ML engineering and data science to build the technical foundation. In 2022, I decided timelines were short enough that it was time to work on AI safety full-time. That's what I've been doing since.

Selected Work

  • Value congruence measurementA framework for testing whether LLMs actually act on their stated values. Models are interviewed about their commitments, then placed in realistic multi-turn scenarios that pressure those commitments — with fresh instances that have no memory of the interview. Measures hypocrisy, not heterodoxy: we don't penalize unusual values, only the gap between what a model says and what it does.
  • Sandbagging detectionResearch on identifying when AI models intentionally perform below their capabilities, including experimental setups across multiple architectures (Llama, Mistral, Phi-3).
  • Biosecurity evaluationsScreening systems, eval frameworks, and dataset curation for assessing biological risks from AI systems (with SecureBio).
  • AI safety educationTeaching and mentoring the next generation of alignment researchers at MATS.

Forecasting

I take AI timelines seriously and have a track record of being early-but-correct on capability milestones. I've won bets on MMLU and MATH benchmark timelines. My current estimate is that things get qualitatively different within the next two years.

Get in touch

I'm looking for collaborators and funders working on technical AI safety, model evaluations, and AI consciousness research.