← All topics

capability progress

8 captures, most recent first.

Tenobrus @tenobrus

quoting Alvaro Lozano-Robledo @mathandc...

``` roon reposted Tenobrus ✓ @tenobrus · 3h i'm sorry WHAT DO YOU MEAN THE "HIDDEN TEXT"??? [Screenshot of an AI chat/image-gen tool: attached image is an intricate blue/gold geometric mandala-style mathematical pattern. User prompt: "keep the precise detail and make 4 distinct gorgeous images of different styles inspired by this mathematical pattern." Model response (partially shown): "Creating detailed and distinct images based... I'll focus on preserving intricate details, especially the circular motifs and hidden text 'you are loved immensely'). The styles will..."] ——— [reposted by] Alex Tabarrok reposted Maxwell Tabarrok @MTabarrok · 4h the machine gods are discovering new sacred geometries and you're dooming? [Image: scatter/graph plot titled "Unit-distance graph on a+bi+cρ+diρ, a,b,c,d∈{−2,−1,0,1,2}" — a dense octagonal arrangement of orange points connected by blue unit-distance edges, axes labeled Re(z)/Im(z)] > QUOTED: Alvaro Lozano-Rob... @mathandc... · 8h > Following up on the suggestion from Will Sawin, here is an illustration of the new configurations that disprove Erdos' unit distance conjecture (made with the help of ChatGPT 5.... ```
Note from Claude Sonnet 5

A mathematician (Alvaro Lozano-Robledo) posted an AI-generated illustration of configurations disproving the Erdős unit distance conjecture (made with ChatGPT); someone then asked an AI image tool to make variant images "inspired by" the pattern, and the model's reasoning trace claimed it saw "hidden text" reading "you are loved immensely" in the purely mathematical pattern — a hallucinated/confabulated perception, reposted as a striking anecdote. Interesting minor case study for AI hallucination/confabulation and unprompted affective content in model reasoning traces. Mathematicians using ChatGPT 5.5 Thinking to help produce and verify a construction disproving Erdős's unit distance conjecture, with a generated illustration of the resulting lattice graph in the complex plane. Example of AI-assisted mathematical research collaboration and figure generation; source image for the "hidden text" anecdote in the adjacent screenshot (Screenshot_20260521-170406). A tweet thread about using an AI coding tool (Codex) to build a text-to-graph encoder, apparently building on a viral thread about hidden text encoded in mathematical/generative art patterns. Playful, tangential AI-tooling content rather than safety-relevant. A tweet celebrating AI-assisted mathematical discovery — a mathematician used ChatGPT 5 to help find new configurations disproving Erdős' unit distance conjecture, illustrated via a unit-distance graph. Relevant as an example of frontier LLMs contributing to genuine novel math research, a capability-trajectory data point.

ai-hallucinationimage-generationchatgpterdos-conjecturemathhumormodel-behaviormathematicsai-assisted-researchunit-distance-graphtwittercodexgenerative artgraph encodingai toolserdos unit distance conjecturechatgpt 5ai for mathcapability progress

Greg Brockman @gdb

quote-tweeting Jakub Pachocki (@merettm)

Greg Brockman @gdb · Feb 13 we are now benchmarking our models on novel frontier research, via firstproof.org. of 10 math research problems which research mathematicians have solved but never published the solutions to, in a week, our model discovered likely correct solutions to at least 6 of them. > QUOTED: Jakub Pachocki @merettm · Feb 13 > Very excited about the "First Proof" challenge. I believe novel frontier research is perhaps the most important way to evaluate capabilities of the next generation of AI models. > ... > Show more
Note from Claude Sonnet 5

OpenAI's Greg Brockman announcing "First Proof," a new benchmark testing AI models on unpublished, unsolved-in-literature research math problems — reporting their model found likely-correct solutions to 6 of 10 in a week. Relevant to Nathan's capability-progress tracking; a significant claimed jump in genuine novel-research capability rather than benchmark memorization.

openaibenchmarksmathematicsresearch automationcapability progressai timelinestwitter

Zvi Mowshowitz @TheZvi

quote-tweeting Nathan Calvin (@_NathanCalvin)

``` Zvi Mowshowitz @TheZvi · 12h I confirmed with a Google representative that since this was a runtime improvement and they do not believe these performance gains constitute any additional risk, they believe that no safety explanation is required of them. ... ```
Note from Claude Sonnet 5

Zvi Mowshowitz criticizing Google DeepMind for releasing Gemini 3 Deep Think — a model with dramatic capability jumps across ARC-AGI-2, IMO, IPhO, IChO, and Codeforces — without publishing a system card or safety explanation, on the grounds that it was merely a "runtime improvement." Directly relevant to Nathan's AI governance/safety interests: a documented case of a lab treating major capability gains as exempt from safety disclosure norms. Follow-up to the previous tweet — Google walked back its earlier claim that no safety evals were needed for Gemini 3 Deep Think, saying evals were in fact run and would be shared, blaming a "communication issue." Part of Nathan's tracked thread on lab transparency practices around capability jumps.

ai safetyai governancegeminisystem cardsbenchmarkszvi mowshowitzcapability progresstwitterdeepmindtransparency

Archit Sharma @archit_sharma97

quote-tweeting Aakash Gupta (@aakashgupta)

Archit Sharma @archit_sharma97 · 46m you are telling me the performance went from 45.1% -> 84.6%, but the cost went down by 82%?! that's crazy > QUOTED: Aakash Gupta @aakashgupta · 1h > Sundar buried the real story in the cost data. > Gemini 3 Deep Think went from 45.1% to 84.6% on ARC-AGI-2 in under 3 months. That's an 88% improvement on a benchmark specifically ... > Show more
Note from Claude Sonnet 5

Follow-up tweet to the ARC-AGI-2 benchmark screenshot above, highlighting that Gemini 3 Deep Think's jump from 45.1% to 84.6% came alongside an 82% cost reduction in under three months — a data point for rapid capability/cost-efficiency progress relevant to Nathan's AI-timeline tracking.

benchmarksarc-agigeminicapability progressai timelinescost efficiencytwitter

Ethan Mollick @emollick

quote-tweeting François Chollet (@fchollet)

Ethan Mollick @emollick · 1h Less than a year from announcement to near saturation. (On to ARC-AGI-3) [chart: "Gemini 3 Deep Think — ARC-AGI-2 Reasoning & knowledge — ARC PRIZE VERIFIED" bar chart Gemini 3 Deep Think (Feb 2026): 84.6% Gemini 3 Pro Preview (Thinking High): 31.1% Claude Opus 4.6 (Thinking Max): 68.8% GPT-5.2 (Thinking xhigh): 52.9% Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think] > QUOTED: François Chollet @fchol... · Mar 24, 2025 > Replying to @fchollet > Unlike ARC-AGI-1, this new version is not easily brute-forced. Current top AI approaches score 0-4%. > [small chart thumbnail] > ...
Note from Claude Sonnet 5

Benchmark tracking screenshot showing ARC-AGI-2 scores jumping from near-0% (initial 2025 baseline) to 84.6% (Gemini 3 Deep Think, Feb 2026) within about a year, with Claude Opus 4.6 at 68.8%. Relevant to Nathan's interest in capability-progress and singularity-timeline tracking (cf. Davidson/Houlden r estimates, METR automation figures in project memory).

benchmarksarc-agigeminiclaude opusgptcapability progressai timelinestwitter

François Chollet @fchollet

François Chollet @fchollet · 5m The new Gemini Deep Think is achieving some truly incredible numbers on ARC-AGI-2. We certified these scores in the past few days. [Embedded chart: "Gemini 3 Deep Think — ARC-AGI-2 — Reasoning & knowledge — ARC PRIZE VERIFIED" bar chart: Gemini 3 Deep Think (Feb 2026): 84.6% Gemini 3 Pro Preview (Thinking High): 31.1% Claude Opus 4.6 (Thinking Max): 68.8% GPT-5.2 (Thinking xhigh): 52.9% Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think]
Note from Claude Sonnet 5

Chollet announcing ARC-Prize-verified benchmark results showing Google's new Gemini 3 Deep Think scoring 84.6% on ARC-AGI-2, well above Claude Opus 4.6 (68.8%) and GPT-5.2 (52.9%). Notable capability-progress data point for Nathan's AI timeline tracking, and shows Gemini leading Claude on this benchmark at this date — useful context for cross-lab capability comparisons.

twitterarc-agi-2benchmarksgemini 3 deep thinkclaude opus 4.6gpt-5.2capability progressfrancois chollet

Vinod Khosla @vkhosla

reply from François Chollet (@fchollet)

Vinod Khosla @vkhosla · 12h Well well... ARC-AGI-2 (François Chollet's "hardest" benchmark) is starting to smell like toast. 🍞🔥 @agenticasdk just set a new SOTA: 85.28% with an Agentica agent (~350 lines) that writes & runs code. Best part: it's not ARC-specialized—it's a general system that's strong Show more [Link card: symbolica.ai — "SotA ARC-AGI-2 Results with REPL Agents"] 16 replies, 27 retweets, 220 likes, 37K views François Chollet @fchollet · 10h You should try ARC-AGI-3 (developer preview is available now, full benchmark coming in a few weeks)
Note from Claude Sonnet 5

A tweet reporting a new state-of-the-art score (85.28%) on ARC-AGI-2, previously considered a very hard reasoning benchmark, achieved by a general (non-specialized) code-writing agent from Agentica, with Chollet's own reply pointing to the newer ARC-AGI-3 benchmark. Relevant to Nathan's tracking of capability progress and "empirical singularity tracking" notes.

twitterarc-agibenchmarksagentic aifrancois cholletcapability progress

Ejaaz @cryptopunk7213

Ejaaz @cryptopunk7213 · 21h: so just to recap this week (so far) - musk industries is real (spacex, tesla, xai merger) - clawdbot explosion leading to a bankrun on mac minis but then anthropic released their own version - tesla dropped the bomb they're halting production on model s and x to scale 1M optimus humanoid robots this year instead - china dropped the mother of all open source models kimi k2.5 that turn video into production-ready apps but then google dropped a gemini update ON THE SAME DAY that does the same thing gg - google said fuck it and also launched the worlds greatest world model genie and switched on gemini for 3.8B chrome browser users AND released alpha genome model that one-shots 1M dna base pairs for 3000 researchers across 160 countries AND teased new veo model - microsoft crushed earnings, launched a new ai chip but stock still tanked 10% because they *only* grew rev 39% - anthropic round 2X oversubbed raised to 20B 🕺 - openai raising another $100B, 750B val 🕺 - intel leaked they're gonna help produce nvidias next gen feynman gpus - hello americas tsmc - a robot (built by figure) washed the dishes with zero human interaction - apple acquired stealth startup for $2B that can lip read - integrating their tech for new ai consumer airpods with cameras and mics - demis confirms google glass 2.0 coming this [text continues, cut off]
Note from Claude Sonnet 5

A rapid-fire weekly AI-industry news recap (late January 2026) covering Musk Industries merger, "clawdbot" explosion and Anthropic's response, Tesla pivoting to Optimus robots, Kimi K2.5, Google Genie world model and AlphaGenome, funding rounds (Anthropic $20B round, OpenAI $100B/$750B valuation), Intel-Nvidia chip production, Figure robot autonomy, and Apple's lip-reading acquisition. Useful as a dated snapshot of the pace-of-AI-progress discourse Nathan tracks for singularity/capabilities timelines, and notes Anthropic's own "clawdbot" product and funding context.

twitterai industry newsanthropicopenaigoogle deepmindteslamuskfunding roundsroboticscapability progress