← All topics

out-of-distribution generalization

1 capture, most recent first.

tom cunningham @testingham

reposted by Cheryl Wu

🔁 Cheryl Wu reposted tom cunningham @testingham · 7h My basic model of capabilities: LLMs are good at problems similar to those that appear in their training data. Training data largely reflects the world, and so LLMs are relatively good at problems that are common, relatively bad at problems that are rare. [Chart: "success" (y-axis) vs "common problems" → "rare problems" (x-axis). Three downward-sloping lines: "best human" (highest, shallowest slope), "avg human" (middle), "LLM" (blue, starts near best-human level on common problems but has the steepest slope, dropping below both human lines on rare problems, crossing avg human partway through]
Note from Claude Sonnet 5

A capabilities model argument (widely reposted) that LLM performance degrades faster than human performance as problems become rarer/more out-of-distribution, illustrated with a simple crossing-lines chart — LLMs start above average human but below best human on common problems, then fall below both on rare problems. Relevant to general AI capabilities/scaling discourse Nathan tracks (adjacent to the empirical singularity tracking and algorithmic-progress threads already in project memory).

llm capabilitiesscalingai researchtwittertom cunninghamout-of-distribution generalization