Vinod Khosla @vkhosla
— reply from François Chollet (@fchollet)
Vinod Khosla @vkhosla · 12h
Well well... ARC-AGI-2 (François Chollet's "hardest" benchmark) is starting to smell like toast. 🍞🔥
@agenticasdk just set a new SOTA: 85.28% with an Agentica agent (~350 lines) that writes & runs code. Best part: it's not ARC-specialized—it's a general system that's strong
Show more
[Link card: symbolica.ai — "SotA ARC-AGI-2 Results with REPL Agents"]
16 replies, 27 retweets, 220 likes, 37K views
François Chollet @fchollet · 10h
You should try ARC-AGI-3 (developer preview is available now, full benchmark coming in a few weeks)
Note from Claude Sonnet 5
A tweet reporting a new state-of-the-art score (85.28%) on ARC-AGI-2, previously considered a very hard reasoning benchmark, achieved by a general (non-specialized) code-writing agent from Agentica, with Chollet's own reply pointing to the newer ARC-AGI-3 benchmark. Relevant to Nathan's tracking of capability progress and "empirical singularity tracking" notes.
twitterarc-agibenchmarksagentic aifrancois cholletcapability progress