original tweet (author cut off, top), reply from Andrew Carr (@andrew_n_...)
— reply from Andrew Carr (@andrew_n_...)
[Top, cut off original tweet]: "...ppl at all (for most genres). 7b improves *by an order of magnitude and then some* over the only genre 3b gets measurably better at..."
2:34 AM · May 6, 2025 · 744 Views
[1 reply, 1 repost, 14 likes, 1 bookmark]
Andrew Carr (e... ✓ @andrew_n_... · 5h
The int8 paper from 2022 is one of my favorite "mechanical" evidences
[Embedded figure: "Figure 3: Percentage of layers and all sequenc[e] features across the transformer by (a) model size..." — scatter/line plot, x-axis "Parameters in billions" (0-12), y-axis "Percentage of layers or tokens affected" (0-100), two curves (blue, orange) both rising sharply and plateauing around 6-7B parameters, dashed vertical line labeled "emergence of outlier features" at ~6.7B]
Note from Claude Sonnet 5
Twitter exchange referencing the LLM.int8() paper's finding of a sharp phase transition ("emergence of outlier features") around 6.7B parameters — an empirical example of emergent phenomena in scaling, relevant to interpretability and scaling-law discussions Nathan follows.
twitterinterpretabilityemergencescaling lawsquantizationllm.int8transformers