← All topics

devin

2 captures, most recent first.

Cognition @cognition

reposted by Jarrod Watts; also @elonmusk (Elon Musk) — saved image

Cognition @cognition · 3h
Grok 4.6 is now available in Devin.

Grok 4.6 marks a significant improvement over Grok 4.5, surpassing GPT-5.6 Sol, behind only Opus 5 and Fable 5.

[Embedded bar chart]
FrontierCode 1.1 Extended  Score
SWE-1.7: 54.3
GPT-5.6 Terra: 55.8
Claude Sonnet 5: 56.2
Grok 4.5: 56.5
GPT-5.5: 56.7
Kimi K3: 58.2
Claude Opus 4.8: 59.6
GPT-5.6 Sol: 60.6
Grok 4.6: 61.3
Claude Opus 5: 63.6
Claude Fable 5: 64.9
Score is a weighted aggregate of rubric items. Solutions that don't pass blocking criteria receive 0.

39 replies, 76 retweets, 1.5K likes, 153K views

Jarrod Watts reposted
Elon Musk @elonmusk
Grok 4.7 will exceed all current models.

That said, Anthropic is a great company and will probably release improved models soon.

However, the SpaceX training corpus is so awesome & unique that I would be shocked if any model is better at real-world engineering than 4.7.

11:24 AM · Aug 12, 2026 · 246.1K Views
Note from Claude Sonnet 5

Cognition (Devin) tweet announcing Grok 4.6 availability with a FrontierCode 1.1 Extended benchmark bar chart ranking models (Claude Fable 5 highest at 64.9, then Claude Opus 5 at 63.6, then Grok 4.6 at 61.3, etc.), followed by an Elon Musk reply predicting Grok 4.7 will exceed all current models due to the 'SpaceX training corpus'.

grokclaude fableclaude opusbenchmarksdevinelon musktwitter

Jared Zoneraich @imjaredz

reposted by Tyler John

``` Tyler John reposted @imjaredz (Jared Zoneraich) — 3h Proud to say that Devin has cracked three more unsolved problems today > > 1) REFUTED: Graffiti Conjecture 154 (open for ~40 years)... > [same three-panel chart image as previous screenshot, showing Graffiti conjecture 154, Graffiti conjectures 39 & 40, and Brandt's regular-supergraph conjecture] ```
Note from Claude Sonnet 5

Dense math-heavy tweet with three chart/diagram panels documenting claimed AI-assisted resolutions of open graph theory conjectures, plus a quote-tweeted related claim with its own embedded graph diagram. Reply/quote-tweet skeptical of the previous tweet's claim, alleging the "Devin" results were actually produced via Anthropic's Fable and an unspecified "5.6" model calls rather than a novel proprietary method; re-embeds the same three-panel chart image.

ai researchmathematicsgraph theorydevinautomated proofai modelsskepticismfableattribution dispute