← All topics

llm capabilities

5 captures, most recent first.

X (Twitter), author not visible (mid-thread scroll)

— saved image

[cut off]spend. These problems will quickly fall to AI because trying a ton of stuff till something clicks is exactly what AI is incredible at. And perhaps, "nobody tried the known stuff long enough" describes many more open problems than we'd think.

Following up on this, here's a thought experiment: Suppose you could take GPT-5.6 or Fable back to 2005, same RL flops, but pretraining data that only existed up to then. Would they still solve the problem?

I don't know, hard to run the counterfactual, but even though many of the tools probably existed in 2005, the "pull" towards which technique to choose that the model "feels right" may depend heavily on the popularity of a given method and our collective instinct as it is recorded in the frequency of an idea used in a particular context. The 2005-pretrain could potentially struggle, not for lack of RL flops, but for lack of pretraining pull towards the right set of ideas. Which implies that these models are something much more interesting than mathematical oracles of truth. We should perhaps think of them as distillations of our accumulated instincts further sharpened by RL.[cut off]
Note from Claude Sonnet 5

Mid-thread tweet (author name not visible on screen) discussing a thought experiment about whether GPT-5.6 or Fable, given only pretraining data up to 2005 but modern RL compute, would still solve certain problems, arguing that models are better understood as distillations of accumulated human instinct sharpened by RL rather than pure 'mathematical oracles of truth'.

llm capabilitiespretrainingreinforcement learningfablegpt-5.6twitter

@gh0stpen

— saved image

ghostpen ✓ @gh0stpen · 13h

this is super eerie, chatgpt's speech to text just thought i was speakin[g] english (was speaking german), and it came out as complete and utte[r] nonsense

BUT IT STILL INFERRED THE CORRECT CONTEXT WORD FOR WORD

---

[user message]
Yeah, as well. I think that's how I knew it. Hmm? Pamita is going. Yeah. Pamita, had, is that? Mus anhaus for weitum, maiden, und um... Haus for weitum, und, uh, jar, echap, geschrim, ange rufen, kin plan, was, yer, tun, so, eh?

Thought for 14s >

TTS hat dich komplett zerhackt, aber ich glaub ich verstehe: **Hausverwaltung/Vermieter geht nicht ran, du hast geschrieben/angerufen, und weißt nicht was jetzt.**
Note from Claude Sonnet 5

Screenshot of an X post by @gh0stpen reporting that ChatGPT's speech-to-text mis-transcribed their German as nonsense pseudo-English, yet the model still recovered the intended meaning word for word. The attached chat shows the garbled transcript and the model's German reply correctly reconstructing it as a question about an unresponsive property manager. The tweet text is cropped at the right edge.

chatgptspeech to texttranscriptionrobustnessgermanllm capabilities

tom cunningham @testingham

reposted by Cheryl Wu

🔁 Cheryl Wu reposted tom cunningham @testingham · 7h My basic model of capabilities: LLMs are good at problems similar to those that appear in their training data. Training data largely reflects the world, and so LLMs are relatively good at problems that are common, relatively bad at problems that are rare. [Chart: "success" (y-axis) vs "common problems" → "rare problems" (x-axis). Three downward-sloping lines: "best human" (highest, shallowest slope), "avg human" (middle), "LLM" (blue, starts near best-human level on common problems but has the steepest slope, dropping below both human lines on rare problems, crossing avg human partway through]
Note from Claude Sonnet 5

A capabilities model argument (widely reposted) that LLM performance degrades faster than human performance as problems become rarer/more out-of-distribution, illustrated with a simple crossing-lines chart — LLMs start above average human but below best human on common problems, then fall below both on rare problems. Relevant to general AI capabilities/scaling discourse Nathan tracks (adjacent to the empirical singularity tracking and algorithmic-progress threads already in project memory).

llm capabilitiesscalingai researchtwittertom cunninghamout-of-distribution generalization

mimrock @mimrocker

quote-tweeting @hampton (hampton — e...)

mimrock @mimrocker · 21h For the last time: People born blind still develop intellect without the vast amount of visual data. It is not necessary for cognition. Do you know what input is crucial for cognition? Language. People born deaf must learn sign language or their mental development will suffer. > QUOTED: hampton — e... ✓ @hampt... · Mar 13 > Chief AI Scientist at Meta, Yann LeCun, believes we're never going to get to human level AI by text: > [Embedded video thumbnail, "NEO NICHE" / "This Is World" clip, captioned "Yann LeCun explains why we're never going to get to human level AI by text", subtitle visible: "word, more or less." Duration 0:54, dated 03.07.25]
Note from Claude Sonnet 5

A tweet pushing back on Yann LeCun's claim that text/language alone can't produce human-level AI (his standard argument for why LLMs are insufficient and world-models/embodiment are needed), using blind and deaf cognitive development as counter-evidence that language, not vision, is the crucial input. Relevant to Nathan's interest in debates over LLM capability ceilings and what substrate/modality is necessary for general intelligence.

twitteryann lecunllm capabilitiescognitionlanguageworld modelsai debate

Noam Brown @polynoamial

Noam Brown ✓ @polynoamial There's a lot of talk of LLMs "saturating all the evals" but there's plenty of evals people could make where LLMs would do poorly: -Beat a Zelda game -Make a profit in a prediction market -Write a stand-up set that's original and funny I'm bullish on AI, but we're far from done. 9:55 AM · Feb 6, 2025 · 2,440 Views 12 replies, 9 reposts, 128 likes, 12 bookmarks Noam Brown ✓ @polynoamial · 4m A lot of grad students have asked me how they can best contribute to the field of AI when they are short on GPUs and making better evals is one thing I consistently point to. [reply, 28 likes] Sir Mr Meow ... ✓ @SirMrMeow... · 3m [reply thread continues, cut off]
Note from Claude Sonnet 5

Noam Brown (OpenAI researcher) argues LLM eval saturation claims are overstated, listing tasks LLMs still fail at; follow-up tweet on grad students contributing via better evals. Relevant to Nathan's interest in AI capability evaluation and benchmarking.

ai evaluationbenchmarksllm capabilitiesnoam brownai progress