Daya Guo @Guodaya
Daya Guo @Guodaya
The 660B R1-Zero and R1 began running after the release of V3, with training taking approximately 2-3 weeks. The R1 model we referred to prior to this time (e.g., in the V3 tech report) was the R1-Lite or the R1-Lite-Zero.
8:36 PM · Feb 3, 2025 · 23K Views
6 replies, 18 reposts, 179 likes, 32 bookmarks
Alex Volkov (Thur... @altr... · 6h
Are those lites... released as well? Any plans to release them? 👀
2 replies, 10 likes, 1.8K views
Daya Guo @Guodaya · 6h
These lite models are currently used only for internal experiments, and there are no plans to open-source them at the moment.
3 replies, 36 likes, 1.9K views
Zephyr @angelusm0rt1s · 6h
Thank you for the amazing work you and...[cut off]
Note from Claude Sonnet 5
A DeepSeek researcher (Daya Guo) clarifies the training timeline and naming history of the DeepSeek R1 / R1-Zero models (660B parameters, ~2-3 weeks training after V3 release), noting "Lite" variants remain internal-only. Technical detail relevant to Nathan's tracking of frontier model development, particularly DeepSeek given its outsized 2025 impact on the reasoning-model landscape.
twitterdeepseekdeepseek-r1daya guomodel trainingreasoning modelsai capabilities