← Timeline

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@nsaphra

— web clipping, 217 words — published 2025-06-23

Thread by @nsaphra

**Naomi Saphra hiring PhD students** @nsaphra [2025-06-23](https://x.com/nsaphra/status/1937280393688010844) Reasoning is about variable binding. It’s not about information retrieval. If a model cannot do variable binding, it is not good at grounded reasoning, and there’s evidence accruing that large scale can make LLMs worse at in-context grounded reasoning. 🧵 --- **Naomi Saphra hiring PhD students** @nsaphra [2025-06-23](https://x.com/nsaphra/status/1937280395243839629) In controlled settings, models switch from grounded in-context learning to generic information retrieval during training. --- **Naomi Saphra hiring PhD students** @nsaphra [2025-06-23](https://x.com/nsaphra/status/1937280396883886573) In actual language models, models move from solving ICL examples with induction heads, which enable algorithmic solutions that are grounded by context, to function vector heads, which retrieve specific known tasks. --- **Naomi Saphra hiring PhD students** @nsaphra [2025-06-23](https://x.com/nsaphra/status/1937280398616404414) After induction heads form and ICL emerges, models get progressively less human-like in their predictions, leading to divergence in larger models. You can't evaluate these models like humans. They are not remotely humanlike, far better at retrieval --- **Naomi Saphra hiring PhD students** @nsaphra [2025-06-23](https://x.com/nsaphra/status/1937280400365126003) So: it doesn’t matter how good LLMs are at the “information retrieval” version of reasoning tasks (eg, describing algos). No variable binding, no grounded reasoning—and rather than helping, raw scale seems to hurt. We can’t be surprised by these results. --- **Aryaman Arora** @aryaman2020 [2025-06-24](https://x.com/aryaman2020/status/1937400838592643480) what about this paper suggesting that context sensitivity is controllable?