← Timeline

Syzygy Research

@syzygyeng on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Syzygy Research @syzygyeng

reposted by Sichu Lu — saved image

↻ Sichu Lu reposted
Syzygy Research @syzygyeng · Aug 3
Today, we're introducing Mach-1 Additive, a 35 billion parameter model that can inference without ever multiplying by a weight. At 1.7 bits per weight, Mach-1 recovers 95% of the performance of the original full precision model, Qwen 3.6 35b, across 12 agentic and reasoning benchmarks, while being 10x smaller.

At 7GB, Mach-1 comfortably fits on consumer laptops with speeds of up to 120 tokens per second, making local inference not just feasible but useful.

Unlike algorithms like BitNet, our approach requires minimal retraining, under 15 GPU hours, making it scalable to massive LLMs. Over the coming weeks, we will be announcing and serving models of up to 3 trillion parameters compressed using our algorithm. For now, you can visit our website to play with Mach-1 directly in your browser, or download our desktop app.

We couldn't be more excited to launch Mach-1. We're looking forward to an energy efficient future for AI, powered by scaled intelligence density.

[embedded chart image, two panels:]
"Time to answer" — bar chart comparing Mach-1 Small, Bonsai 27B, Gemma 4 Q2 across prompt lengths (128-tok: 3.7s/6.1/11.5; 2,048-tok: 5.5s/9.8/16.0; 8,192-tok: 12.9s/22.3/36.8)
"Speed of intelligence" — bar charts of "Intelligence per second (pts/s)": Mach-1 Small 15.6, Bonsai 27B 9.2, Gemma 4 Q2 4.7; "Intelligence density per second (pts/GB/s)": Mach-1 Small 1.97, Bonsai 27B 1.28, Gemma 4 Q2 0.40
"Methodology" [link text, cut off at bottom]
Note from Claude Sonnet 5

Tweet by Syzygy Research announcing "Mach-1 Additive," a 35B-parameter multiplication-free quantized (1.7 bits/weight) LLM derived from Qwen 3.6 35B, claiming 95% performance retention at 10x smaller size and fast local inference; includes embedded bar charts comparing latency and "intelligence density" against Bonsai 27B and Gemma 4 Q2.

llm quantizationmodel compressionlocal inferencebenchmarks