← All topics

ai safety research

2 captures, most recent first.

david rein @idavidrein

david rein ✓ @idavidrein · 16h To evaluate our monitoring systems, we typically rely on a pretty strong assumption that we've given agents a concrete, narrowly scoped task. But we're going to give agents increasingly abstract/high-level tasks as they get better. This means (to the extent they're misaligned) they'll be able to justify a bunch of pretty unrelated stuff as being part of their assigned goal. I think this might be an important way that our monitorability trends don't capture the actual dynamics here, since we typically evaluate on a fixed distribution of tasks. > QUOTED: david rein @idavidrein · May 9: Frontier AI labs like OpenAI and Anthropic have started using automated LLM monitoring systems to make sure their internal AI agents stay aligned [1, 2]. I think this is really important and good! Unfortunately, it's difficult to get … [truncated by platform]
Note from Claude Sonnet 5

Quote-tweet structure; the quoted tweet text is cut off with an ellipsis indicating platform truncation, not illegibility.

ai alignmentmonitoringagentic aiai safety research

mark bissell @MarkMBissell

mark bissell @MarkMBissell · 19h working in interp means reminding yourself every single day of the first principle > QUOTED: mark biss... @MarkMBis... · Oct 16, 2025 > Replying to @dnbt777 and @RichardMCNgo > the first principle > > [Attached photo of Richard Feynman at a chalkboard, with caption overlay: "The first principle is that you must not fool yourself — and you are the easiest person to fool." — Richard Feynman]
Note from Claude Sonnet 5

An interpretability researcher's tweet invoking Feynman's "don't fool yourself" principle as a daily discipline for interp work — relevant to Nathan's epistemic protocol around not over-interpreting model self-reports and being wary of confident-quickly conclusions in interpretability/consciousness research.

interpretabilityfeynmanepistemicsai safety researchtwitter