— web clipping, 393 words — published 2026-03-31
Thread by @TheOneKloud
**Pierre Richemond** @TheOneKloud 2026-03-31
1/8:Huge fan of @SchmidhuberAI's foundational work. Since BYOL is being discussed, I wanted to share some thoughts as a co-author who helped shape its intellectual foundations on where it comes from and why it works — which I think involves a confluence of three distinct pillars. 🧵
> 2026-03-31
>
> Dr. LeCun's heavily promoted Joint Embedding Predictive Architecture (JEPA, 2022) \[5\] is the heart of his new company. However, the core ideas are not original to LeCun. Instead, JEPA is essentially identical to our 1992 Predictability Maximization system (PMAX) \[1\]\[14\].
>
> [image]
---
**Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039084626359066998)
2/8: Pillar 1: The latent prediction objective. PMAX ('92) explored non-generative prediction of abstract representations, building on IMAX (Becker & Hinton, '89). PMAX's specific contribution was exploring explicit anti-collapse mechanisms (like variance penalties).
---
**Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085446076969074)
3/8: Pillar 2: The co-training principle (Blum & Mitchell, '98)—the insight that cross-view prediction from conditionally independent views extracts robust semantics. This was a recurring topic of conversation with @geoffreyhinton right after BYOL. I cited it in my PhD thesis as
[image]
---
**Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085600855187922)
4/8: Pillar 3: Implicit optimization dynamics. This is the crux of what makes BYOL different. PMAX noted prediction and collapse-avoidance are conflicting goals, solving it with explicit penalties. BYOL's discovery: with the right architecture, the explicit penalty can be zero.
---
**Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085766228242536)
5/8: So why doesn't BYOL collapse? In our ICML 2023 paper "The Edge of Orthogonality," we proved that the predictor + EMA + stop-gradient implicitly enforce quasi-orthogonality, making BYOL equivalent to Riemannian gradient descent on a Stiefel manifold.
---
**Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085896390062303)
6/8: The "negative terms" needed to prevent collapse aren't in the loss—they emerge from the geometry of staying close to the orthogonal manifold (Prop. 9). We validated this constructively: our closed-form orthonormalizing predictors outperform the learned one on ImageNet.
---
**Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039086083074330736)
7/8: Our predictors use Newton-Schulz iterations for fast approximate orthonormalization. Great to see this mathematical primitive gain wider traction since, notably in the Muon optimizer. Orthogonalization may be a deeper principle than any single application of it.
---
**Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039086354181578938)
8/8: BYOL = latent prediction + co-training + implicit architectural orthogonalization. Science is cumulative, and proper credit—in all directions—matters. Paper:
---
**Prosperous727** @prosperous727 [2026-04-01](https://x.com/prosperous727/status/2039173118246170957)
@Grok how can The Architect-1 benefit this?