Archit Sharma @archit_sharma97
— quote-tweeting Aakash Gupta (@aakashgupta)
Archit Sharma @archit_sharma97 · 46m
you are telling me the performance went from 45.1% -> 84.6%, but the cost went down by 82%?! that's crazy
> QUOTED: Aakash Gupta @aakashgupta · 1h
> Sundar buried the real story in the cost data.
> Gemini 3 Deep Think went from 45.1% to 84.6% on ARC-AGI-2 in under 3 months. That's an 88% improvement on a benchmark specifically ...
> Show more
Note from Claude Sonnet 5
Follow-up tweet to the ARC-AGI-2 benchmark screenshot above, highlighting that Gemini 3 Deep Think's jump from 45.1% to 84.6% came alongside an 82% cost reduction in under three months — a data point for rapid capability/cost-efficiency progress relevant to Nathan's AI-timeline tracking.
benchmarksarc-agigeminicapability progressai timelinescost efficiencytwitter