← Timeline

Harry Partridge

@part_harry_ on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Harry Partridge @part_harry_

quoting Charlie O'Neill (@oneill_c)

Harry Partridge @part_harry_ Pretraining is data-inefficient. This is entirely a consequence of the fact that we throw away the KV cache after every forward-backward step! If we can integrate efficient KV cache compaction into pretraining, we will unlock human level data efficiency. Neural KV cache compaction makes this possible. [quoted tweet:] Charlie O'Neill @oneill_c · 2h [image: 8-panel chart grid — Train: Loss, Train: Utilization, Eval: MCQ Accuracy, Eval: MCQ Utilization, Eval: CE Utilization, Eval: Cont Utilization, Eval: Compact KL, Eval: Compact CE — all plotted against Number of Latents] "Towards infinite context windows: neural KV cache compaction" TL;DR: We compress an LLM's KV cache 8x in milliseconds, retaining 85%+ of factual accuracy across domains. Here's how. Introduction... 10:53 AM · Apr 1, 2026 · 3,243 Views [reply] 1 [retweet] 5 [heart] 45 [bookmark] 34 [share] Harry Partridge @part_harry_ · 1h In context learning *is* data efficient. We spend so much compute building this rich KV cache on every forward pass. If we can effectively merge this cache with our existing knowledge, we can dramatically improve the data efficiency of pre-training and RL!
Note from Claude Sonnet 5

Twitter thread on ML research about neural KV cache compaction for infinite context windows, and a hypothesis linking it to pretraining data efficiency. Technical ML architecture content relevant to Nathan's brain_graph_1 / memory-architecture interests.

machine learningkv cachetransformerspretrainingdata efficiencytwitterml research