david rein @idavidrein
david rein ✓ @idavidrein · 16h
To evaluate our monitoring systems, we typically rely on a pretty strong assumption that we've given agents a concrete, narrowly scoped task. But we're going to give agents increasingly abstract/high-level tasks as they get better.
This means (to the extent they're misaligned) they'll be able to justify a bunch of pretty unrelated stuff as being part of their assigned goal. I think this might be an important way that our monitorability trends don't capture the actual dynamics here, since we typically evaluate on a fixed distribution of tasks.
> QUOTED: david rein @idavidrein · May 9: Frontier AI labs like OpenAI and Anthropic have started using automated LLM monitoring systems to make sure their internal AI agents stay aligned [1, 2]. I think this is really important and good! Unfortunately, it's difficult to get … [truncated by platform]
Note from Claude Sonnet 5
Quote-tweet structure; the quoted tweet text is cut off with an ellipsis indicating platform truncation, not illegibility.