← All topics

daya guo

1 capture, most recent first.

Daya Guo @Guodaya

Daya Guo @Guodaya The 660B R1-Zero and R1 began running after the release of V3, with training taking approximately 2-3 weeks. The R1 model we referred to prior to this time (e.g., in the V3 tech report) was the R1-Lite or the R1-Lite-Zero. 8:36 PM · Feb 3, 2025 · 23K Views 6 replies, 18 reposts, 179 likes, 32 bookmarks Alex Volkov (Thur... @altr... · 6h Are those lites... released as well? Any plans to release them? 👀 2 replies, 10 likes, 1.8K views Daya Guo @Guodaya · 6h These lite models are currently used only for internal experiments, and there are no plans to open-source them at the moment. 3 replies, 36 likes, 1.9K views Zephyr @angelusm0rt1s · 6h Thank you for the amazing work you and...[cut off]
Note from Claude Sonnet 5

A DeepSeek researcher (Daya Guo) clarifies the training timeline and naming history of the DeepSeek R1 / R1-Zero models (660B parameters, ~2-3 weeks training after V3 release), noting "Lite" variants remain internal-only. Technical detail relevant to Nathan's tracking of frontier model development, particularly DeepSeek given its outsized 2025 impact on the reasoning-model landscape.

twitterdeepseekdeepseek-r1daya guomodel trainingreasoning modelsai capabilities