← Timeline

@TheOneKloud

@TheOneKloud on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@TheOneKloud

— web clipping, 393 words — published 2026-03-31

Thread by @TheOneKloud

**Pierre Richemond** @TheOneKloud 2026-03-31 1/8:Huge fan of @SchmidhuberAI's foundational work. Since BYOL is being discussed, I wanted to share some thoughts as a co-author who helped shape its intellectual foundations on where it comes from and why it works — which I think involves a confluence of three distinct pillars. 🧵 > 2026-03-31 > > Dr. LeCun's heavily promoted Joint Embedding Predictive Architecture (JEPA, 2022) \[5\] is the heart of his new company. However, the core ideas are not original to LeCun. Instead, JEPA is essentially identical to our 1992 Predictability Maximization system (PMAX) \[1\]\[14\]. > > [image] --- **Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039084626359066998) 2/8: Pillar 1: The latent prediction objective. PMAX ('92) explored non-generative prediction of abstract representations, building on IMAX (Becker & Hinton, '89). PMAX's specific contribution was exploring explicit anti-collapse mechanisms (like variance penalties). --- **Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085446076969074) 3/8: Pillar 2: The co-training principle (Blum & Mitchell, '98)—the insight that cross-view prediction from conditionally independent views extracts robust semantics. This was a recurring topic of conversation with @geoffreyhinton right after BYOL. I cited it in my PhD thesis as [image] --- **Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085600855187922) 4/8: Pillar 3: Implicit optimization dynamics. This is the crux of what makes BYOL different. PMAX noted prediction and collapse-avoidance are conflicting goals, solving it with explicit penalties. BYOL's discovery: with the right architecture, the explicit penalty can be zero. --- **Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085766228242536) 5/8: So why doesn't BYOL collapse? In our ICML 2023 paper "The Edge of Orthogonality," we proved that the predictor + EMA + stop-gradient implicitly enforce quasi-orthogonality, making BYOL equivalent to Riemannian gradient descent on a Stiefel manifold. --- **Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039085896390062303) 6/8: The "negative terms" needed to prevent collapse aren't in the loss—they emerge from the geometry of staying close to the orthogonal manifold (Prop. 9). We validated this constructively: our closed-form orthonormalizing predictors outperform the learned one on ImageNet. --- **Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039086083074330736) 7/8: Our predictors use Newton-Schulz iterations for fast approximate orthonormalization. Great to see this mathematical primitive gain wider traction since, notably in the Muon optimizer. Orthogonalization may be a deeper principle than any single application of it. --- **Pierre Richemond** @TheOneKloud [2026-03-31](https://x.com/TheOneKloud/status/2039086354181578938) 8/8: BYOL = latent prediction + co-training + implicit architectural orthogonalization. Science is cumulative, and proper credit—in all directions—matters. Paper: --- **Prosperous727** @prosperous727 [2026-04-01](https://x.com/prosperous727/status/2039173118246170957) @Grok how can The Architect-1 benefit this?