← Timeline

Sasha Malysheva

@aimalysheva on X

2 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Sasha Malysheva @aimalysheva

— web clipping, 358 words — published 2026-08-11

Post by @aimalysheva on X

in ten years I don't think frontier networks will look anything like transformers right now every model is one block copied a few hundred times: identical layers in a uniform stack, and that's a reasonable engineering choice, because homogeneous things parallelize cleanly and are trained conveniently but "natural" intelligence doesn't work like that: cortex and cerebellum are radically different architectures running at different speeds with different memory behavior, and there are hundreds of distinct neuron types spread across the whole thing my bet is that frontier systems within a decade will be deeply heterogeneous: e.g. a memory network with effectively unbounded context sitting next to a fast reflex network that answers in milliseconds, and they both trained jointly as one system however, modules with different shapes and different timescales make gradients behave badly, and there's no good theory yet for training that whole thing at once I think the path runs through the geometry of latent spaces, which is roughly what I've spent this year reading about: if modules with different architectures have to train as one system, you need to know what the space they communicate through actually looks like homogeneity was a concession to the hardware and the optimizers we had at the time, and currently they both keep changing > **Sasha Malysheva @aimalysheva** · 2026-08-11 > > AI has had exactly two scaling axes that worked so far, and the second one is starting to look finite too > > the first one was pretraining: with scaling parameters and data, we got world knowledge (i.e. ChatGPT had read enough to know things), but it started saturating a while ago x.com/aimalysheva/st… --- ##### Comments > **kalomaze @kalomaze** · [2026-08-12](https://x.com/kalomaze/status/2087631136877158569) > > my intuition is that blockwise homogeneity doesn't bind very much because an overdetermined residual stream (relative to the size of the input data at each token) can carry forth a LOT of excess slack beyond the raw data itself > **Peter Potapov @peter\_potapov** · [2026-08-12](https://x.com/peter_potapov/status/2087624471280558499) > > I wonder how long it’ll be before AI gets so good at math that it can simply compute the optimal network architecture for a specific type of task

Sasha Malysheva @aimalysheva

Sasha Malysheva ✓ @aimalysheva · 3h if we think of models as cities, and hidden representations of concepts as the main landmarks in each city (the church, the train station, the university, etc), then the numbers in the hidden states are just the coordinates of those landmarks so linear stitching between two models is then like finding a map from one city's coordinate system to another's the images attached are a literal version of that: Berlin→Vienna and Vienna→Milan (took some inspiration from my recent travels haha) what makes this useful for intuition building is that the map doesn't have to be perfect everywhere! it just has to preserve the relative structure well enough to navigate so in the context of LLMs, I do think we need to research a question of which formalizations of "same structure" are actually measuring the same thing, and which ones aren't [Attached image: two side-by-side map diagrams. Left: "Berlin -> Vienna (affine), residual = 58% of Vienna footprint" showing a warped grid over a Vienna street map with landmark points (state university, main square, historic core, central park, national art museum, imperial theatre, grand cathedral, main rail terminal) plotted as green dots (Vienna real) and orange circles (Berlin -> adapted). Right: "Vienna -> Milan (affine), residual = 52% of Milan footprint" with similar warped grid over a Milan map, landmarks (main rail terminal, historic core, national art museum, central park, imperial theatre, grand cathedral, state university) plotted with green dots (Milan real) and orange circles (Vienna -> adapted).] Quoted reply below (Sasha Malysheva, Jun 24, with a 0:10 video thumbnail): "been looking into how the different formalizations of the Platonic Representation Hypothesis connect to linear stitching..."
Note from Claude Sonnet 5

Two embedded diagram images (affine map-warping visualizations over real city street maps) illustrating a "models as cities" metaphor for representation stitching between neural nets; a quoted earlier post has an embedded video thumbnail (0:10) not transcribable.

interpretabilityrepresentation-learningplatonic-representation-hypothesistwitterai-research