← All topics

rysana

1 capture, most recent first.

X (Twitter) — unclear handle (bottom of thread cut off), replying to @jrysana (John) and @sdmat123

replying to @jrysana (John) and @sdmat123

[top of tweet cut off above frame] 700,000 tokens per second on one request with zero caches etc at near perfect score never gets old. What. The. Fuck. It is so fast that it's as if it didn't even generate, it just already existed somehow. Yet it didn't exist – it's brand new, from the prompt sent 10ms ago. 10:51 PM · Jun 27, 2026 · 813 Views 💬3 🔁 ♥19 🔖 ⤴ Relevant ⌄ John (@jrysana) — 9h As far as I know, these are roughly the next-best systems out there: Special chip, peak ~35,000 tok/sec but very small very quantized model, unfortunately just not usable yet Special chip, peak ~6000 tok/sec on very small dumb models - and apparently not much scalability 💬 🔁 ♥4 📊133 🔖 ⤴ John (@jrysana) — 9h 900,000* 💬 🔁 ♥1 📊142 🔖 ⤴ sdmat (@sdmat123) — 3h What's this?! 💬1 🔁 ♥2 📊26 🔖 ⤴ John (@jrysana) — 3h Rysana V2, up to ~10,000x more efficient than current frontier models, recent result (from yesterday)
Note from Claude Sonnet 5

Thread about extreme LLM inference speed (claimed 700,000–900,000 tokens/sec, "Rysana V2"); top tweet's author handle and opening lines are cut off above the visible frame.

llm inference speedai hardwarerysanatwitterbenchmarks