← All topics

ai hardware

2 captures, most recent first.

Lisan al Gaib @scaling01

— saved image

Lisan al Gaib ✅ @scaling01 · 30m
Cerebras is talking about 10T models running at 1000 tokens/s

[embedded chart, titled "CS-4 ENABLES SUB-1MS LATENCY (1000 TOK/S) FOR 10T MODELS AND BEYOND"]
Chart: "WAFER-TO-WAFER LATENCY VS. MODEL SIZE" — line graph, x-axis "Model Size (Trillion of Parameters)" 1-10, y-axis "IO latency across all hops (ms)" 0.0-0.6. CS-3 (purple line) rises from ~0.07ms to ~0.58ms; CS-4 (orange line) rises from ~0.02ms to ~0.2ms. Annotations: "2.5X FASTER / 2.5X LOWER LATENCY", "0.2MS LATENCY ACROSS ALL HOPS FOR 10T PARAMETERS". Source: Internal benchmarking and projection (August 2026).

[quoted tweet]
Lisan al Gaib ✅ @scaling01 · 34m
[small chart thumbnail comparing CS-3 vs CS-4 specs]
Cerebras just announced their new AI accelerator CS-4
Note from Claude Sonnet 5

Tweet about Cerebras announcing its CS-4 AI accelerator chip, claiming sub-1ms latency and 1000 tokens/s for 10-trillion-parameter models, with a company benchmark chart comparing CS-3 vs CS-4 wafer-to-wafer latency scaling.

cerebrasai hardwarescalingtwitter

X (Twitter) — unclear handle (bottom of thread cut off), replying to @jrysana (John) and @sdmat123

replying to @jrysana (John) and @sdmat123

[top of tweet cut off above frame] 700,000 tokens per second on one request with zero caches etc at near perfect score never gets old. What. The. Fuck. It is so fast that it's as if it didn't even generate, it just already existed somehow. Yet it didn't exist – it's brand new, from the prompt sent 10ms ago. 10:51 PM · Jun 27, 2026 · 813 Views 💬3 🔁 ♥19 🔖 ⤴ Relevant ⌄ John (@jrysana) — 9h As far as I know, these are roughly the next-best systems out there: Special chip, peak ~35,000 tok/sec but very small very quantized model, unfortunately just not usable yet Special chip, peak ~6000 tok/sec on very small dumb models - and apparently not much scalability 💬 🔁 ♥4 📊133 🔖 ⤴ John (@jrysana) — 9h 900,000* 💬 🔁 ♥1 📊142 🔖 ⤴ sdmat (@sdmat123) — 3h What's this?! 💬1 🔁 ♥2 📊26 🔖 ⤴ John (@jrysana) — 3h Rysana V2, up to ~10,000x more efficient than current frontier models, recent result (from yesterday)
Note from Claude Sonnet 5

Thread about extreme LLM inference speed (claimed 700,000–900,000 tokens/sec, "Rysana V2"); top tweet's author handle and opening lines are cut off above the visible frame.

llm inference speedai hardwarerysanatwitterbenchmarks