← All topics

cognition ai

1 capture, most recent first.

@silasalberti

quoting @cognition

Silas Alberti ✔ [icon] @silasalberti · 6h Recently the industry chatter moved from RL to pretraining (& midtraining): RL is supposed to get diminishing returns because it hits a "ceiling" that depends on the quality of your pretrain. While that is certainly true, it's an open question how high the ceiling actually is. We wanted to push this to the limit: for Kimi K2.7 which has already been heavily post-trained multiple times, it initially seemed like there wasn't much to squeeze. We dealt with fast entropy collapse & many other issues. However, we used this as a challenge to tune our algorithm and data. At some point, it suddenly started working and the model, SWE-1.7, magically improved by much more than we expected. It also seems like the ceiling isn't obviously hit yet, so we're excited to see how much further we can push it. [Quoted tweet:] Cognition ✔ @cognition · 7h Introducing SWE-1.7, the most capable model we've trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is ... [Embedded table comparing benchmark scores: SWE-1.7, GLM-5.2, Composer 2.5, Opus 4.8, GPT-5.5 — FrontierCode 1.1 Main: 42.3%, 24.5%, 25.6%, 46.5%, 43.0% Terminal-Bench 2.1: 81.5%, 81.0%, 76.0%, 86.9%, 84.2% SWE-Bench Multilingual: 77.8%, 74.5%, 71.6%, 84.4%, 76.8%]
Note from Claude Sonnet 5

Technical AI-research tweet discussing reinforcement-learning post-training results on the Kimi K2.7 base model (yielding "SWE-1.7"), with an embedded benchmark comparison table against GLM-5.2, Composer 2.5, Opus 4.8, and GPT-5.5.

llm trainingreinforcement learningbenchmarkscognition aitwitter