← All topics

ai-capabilities

6 captures, most recent first.

François Chollet @fchollet

François Chollet ✓ @fchollet · 1h Most human tasks are not Markovian, the optimal next action cannot be determined solely by looking at the current state. It depends heavily on the past trajectory, the original intent, and context constraints. An agent that cannot compress and track its past trajectory with absolute fidelity is maybe 20% as useful as one that can.
Note from Claude Sonnet 5

Chollet argument about agent memory/context-tracking fidelity as a bottleneck for agentic usefulness, since most real tasks are non-Markovian and depend on trajectory history rather than current state alone. Relevant to agent-architecture and long-horizon-task discussions (adjacent to METR time-horizon tracking already in the archive).

agentsmemorycontextmarkovianai-capabilitieschollet

Noam Brown @polynoamial

Noam Brown @polynoamial · Jan 26 1987: AI can't win at chess—planning is uniquely human 1997: AI can't win at Go—intuition is uniquely human 2016: AI can't win at poker—bluffing is uniquely human 2023: AI can't get IMO gold—reasoning is uniquely human 2026: AI can't make wise decisions—judgment is uniquely human [Screenshot of NYT-style opinion guest essay, headline partially visible: "OPINION GUEST ESSAY ... [Hu]mans Poss[ess a] Thing Tha[t AI Does] Not: Judg[ment]" — visible body text fragment: "...hean by 'judgment'? The...mong competing values a[re a matter]...of opinion, to weigh consi[derations]...independently but canno[t weigh them]...at once, to consider seve[ral]...ght on the best one. Judg[ment]...ely on when trade-offs ar[ise]...e and the right answer is [contested]...uted. It is a uniquely hum[an capacity]" (last clause highlighted in blue)]
Note from Claude Sonnet 5

Noam Brown (OpenAI researcher, known for poker/Diplomacy AI) mocking a recurring pattern of "AI can't do X, X is uniquely human" claims that keep getting falsified, applied here to a 2026 NYT opinion essay claiming judgment/wisdom is the next uniquely-human bastion. Relevant to capability-timeline tracking and the recurring rhetorical pattern of moving goalposts on AI capability claims.

ai-capabilitiesforecastingnoam-browntwitterjudgmentnyt-opiniongoalpost-moving

Nabeel S. Qureshi @nabeelqu

quoting/referencing Neel Somani (@neelsomani); scroll also shows Techartist (@techartist_)

[Top, cut off]: "@AnthropicAI please sponsor this man." 💬76 🔁64 ❤️1.1K 📊259K views [bookmark] [share] Nabeel S. Qureshi ✓ @nabeelqu · Jan 18 He put an Erdos Problem into GPT 5.2Pro, no special prompting, and it *thought continuously for 41 minutes and spat out a correct proof*. It's over everyone, we had a good run. [Embedded screenshot of a GPT reasoning session]: "Let $n_1<n_2<\cdots$ be an infinite sequence such that, for any choice of congruence classes $a_i \pmod{n_i}$, the set of integers not satisfying any of the congruences $a_i \pmod{n_i}$ has density $0$. Is it true that for every $\epsilon>0$ there exists some $k$ such that, for every choice of congruence classes $a_i$, the density of integers not satisfying any of the congruences $a_i \pmod{n_i}$ for $1\leq i\leq k$ is less than $\epsilon$?" ↳ Reasoned for 41m 54s "I'm parsing the problem carefully. We have an infinite sequence of moduli and residue classes, and the integers not satisfying any congruence are of density 0. This suggests that a union of these arithmetic progressions should cover almost all integers. But does it imply that a finite subcollection might already cover most of the integers for all residue choices? I need to explore if this resembles compactness or something like Borel-Cantelli. I'm considering the infinite property of moduli and how they form a "covering system," meaning they cover almost all integers no matter the residues chosen. The complement of covered integers, U(a), has zero density. Is this uniform across all sets of residues? I'm thinking this might relate to Dini's theorem, but could the uncovered density decrease slowly..." 281 Anonymous · Jan 16, 2026 [Open in browser] Neel Somani ✓ @neelsomani · Jan 10 [Embedded image: math proof example with equations, "Example (a=2). Here c=49, and the identity becomes..."] Weekend win: The proof I submitted for Erdos Problem #397 was accepted by Terence Tao. The proof was generated by GPT 5.... 💬47 🔁103 ❤️1.4K 📊215K Techartist ✓ @techartist_ · 23h Interactive quantum neural network built with Three.js and GLSL shaders, wrapped in a glassmorphic UI. Click or drag sends pulses while f[orm, colors, and density update in real time through...] [cut off]
Note from Claude Sonnet 5

Twitter scroll showing GPT-5.2 Pro reportedly solving an open Erdos problem after 41+ minutes of extended reasoning, plus a related tweet about a GPT-5-generated proof for Erdos Problem #397 accepted by Terence Tao. Directly relevant to Nathan's tracking of frontier-model mathematical capability and empirical singularity/AI-R&D-automation signals.

gpt-5mathematical-reasoningerdos-problemsterence-taoai-capabilitiestwitterextended-thinkingagi-progress

Sauers @Sauers_

Sauers ✓ @Sauers_ · Jan 10 People who think Opus 4.5 is best at everything probably have less complex problems. Codex 5.2 and Gemini 3 Pro have unique and valuable forms of intelligence [Illustration: three abstract stylized "brain" icons side by side — an orange swirling organic brain shape, a blue faceted crystalline/circuit brain shape, and a black cloud-like brain shape filled with a galaxy/starfield.]
Note from Claude Sonnet 5

A tweet arguing against the idea that one frontier model (Opus 4.5) dominates across all task types, asserting that Codex 5.2 and Gemini 3 Pro have distinct cognitive strengths. Directly relevant to Nathan's ongoing model-individuation interest — different model "topologies" rather than a single intelligence ranking — though from an outside commentator rather than a model self-report.

twittermodel-comparisonmodel-individuationopus-4.5codexgemini-3-proai-capabilities

Nathan Labenz @labenz

Nathan Labenz @labenz · 1h AIs are weird, but I'm using OpenAI Operator daily right now Here it's going to AI Studio where Gemini 2.5 Pro is answering questions based on 500K+ tokens context [Screenshot of Operator agent interaction] I am unable to copy the content from the Google AI Studio page, as the page does not allow me to copy the generated text. Would you like me to summarize the content and provide it in a different format? [User reply, gray bubble] hmm, it does allow me to copy & paste – but ... you can also just copy and paste "mentally" – you can write, right? so just ... mentally copy & paste if needed?
Note from Claude Sonnet 5

Nathan Labenz tweet showing an amusing agentic-AI failure: OpenAI's Operator agent claims it can't copy text from a webpage, and Labenz sarcastically instructs it to "mentally" copy-paste. Illustrates agent tool-use limitations/confabulated constraints in a computer-use agent, relevant to AI capabilities/agent-reliability observations.

twitternathan-labenzopenai-operatoragentic-aigeminicomputer-useai-capabilities

Riley Goodside @goodside

Riley Goodside @goodside · 6h The Screenshot: a fake screenshot generated by ChatGPT 4o of a Wikipedia article about the screenshot itself, with a copy of the screenshot in the article [AI-generated fake Wikipedia page, garbled/glitchy text throughout, recursive imagery:] WIKIPEDIA The Free Encyclopenia Talk | Contributions | Creats account | Search The Screenshot (self-referential image) [infobox] This article is about the self-referential image generated by ChatGPT. For information about screenshots in general, see Screenshot. The Screenshot (self-referential image) [embedded smaller copy of the same fake Wikipedia screenshot, recursively nested, increasingly illegible/garbled text at smaller scale] Description – The Screenshot" is a image depicting a screenshot of a Wikipedia article titled "The Screenshot" that describes the image itself. The image is generated by ChatGPT 4o, a large language model developed by OpanAI. The image is recursive in nature, showing the Wikipedia interface includin the logo, navigation menu, and other standard elements of a Wikipedia article, with a lead section that contextualizes image for the reader. See also
Note from Claude Sonnet 5

Riley Goodside (prominent AI/prompting researcher) shares a ChatGPT-4o-generated image that recursively depicts a Wikipedia article about itself, containing a nested copy of itself — a demonstration of image-generation self-reference and the model's characteristic text-rendering glitches/garbling at small scale. Interesting artifact of AI image-gen capability/limitation from a well-known AI researcher's account.

twitterriley-goodsidegpt-4oimage-generationself-referenceai-capabilitieswikipedia-parody