← All topics

autonomy horizon

1 capture, most recent first.

Alex Ratner @ajratner

Alex Ratner (verified) @ajratner · 3h There are three major vectors of progress for AI capabilities, and the benchmarks that measure them: (1) Environment complexity --> E.g. complex, domain-specific context and tool/action spaces, human interaction, world modeling (2) Autonomy horizon --> E.g. long horizon, non-stationary goals (3) Output complexity --> E.g. complex outputs with nuanced, rubric-based evaluation / reward signals We are just beginning to systematically *measure* tasks with truly complex inputs and envs, complex outputs/rubrics, and long horizon execution - let alone solve them. The frontier remains open!
Note from Claude Sonnet 5

A framework from Snorkel AI's Alex Ratner categorizing three axes of AI capability progress (environment complexity, autonomy horizon, output complexity) and their corresponding benchmarks. Relevant to Nathan's tracking of empirical AI capability/singularity signals (autonomy horizon connects directly to METR-style task-length measurements referenced elsewhere in the archive).

twitterai capabilitiesbenchmarksautonomy horizonevaluationagi timelines