← Timeline

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Sakura Yuki @sakurayukiai

Sakura Yuki @sakurayukiai A 97% drop in optimizer VRAM during RL training sounds like a stunt, but the math is so clean. Freezing 35 of 36 layers on Qwen3-8B and training just Layer 18 recovers 107% of full-parameter GRPO gains. Reasoning is purely a middle-layer story?? 10:03 AM · Jul 2, 2026 · 15.7K Views [11 replies, 10 reposts, 206 likes, 143 bookmarks]
Note from Claude Sonnet 5

Plain text tweet; avatar photo of a woman in sunglasses.

twittermachine-learningrl-trainingqwentechnical