← All topics

distributed training

1 capture, most recent first.

Simo Ryu @cloneofsimo

reply from Pavel Surmenok (@surmenok)

Simo Ryu @cloneofsimo Many noise from recent 4.5 release, but two important take from vid imo: 1. "GPT4.5 was trained on multiple datacenters" Translate that to "diloco goes brr for largest LLM on the market", bullish on async, low bandwidth training in 2025. 2. "We aggressively used low precision training" -> another use of fp8 training, presumably on h100s. Im guessing they benefited from fp8 because of high granularity 3:44 AM · Feb 28, 2025 · 10.9K Views [4 replies, 4 reposts, 120 likes, 25 bookmarks] Pavel Surmenok @surmenok · 29m It doesn't have to be async. Google is training on multiple datacenters synchronously. You just need a high bandwidth link. [40 views shown] subho ghosh @SubhoGhosh02 · 11h (partially obscured by nav bar) deepseek is way ahead :)
Note from Claude Sonnet 5

A technical Twitter thread analyzing GPT-4.5's training details (multi-datacenter training implying DiLoCo-style distributed/async training, aggressive fp8 low-precision training) with a corrective reply noting Google trains synchronously across datacenters via high-bandwidth links. Relevant to Nathan's tracking of frontier-lab training infrastructure and scaling techniques.

twittergpt-4.5distributed trainingfp8dilocoscalingtraining infrastructure