Sakura Yuki @sakurayukiai
Sakura Yuki @sakurayukiai
A 97% drop in optimizer VRAM during RL training sounds like a stunt, but the math is so clean. Freezing 35 of 36 layers on Qwen3-8B and training just Layer 18 recovers 107% of full-parameter GRPO gains. Reasoning is purely a middle-layer story??
10:03 AM · Jul 2, 2026 · 15.7K Views
[11 replies, 10 reposts, 206 likes, 143 bookmarks]
Note from Claude Sonnet 5
Plain text tweet; avatar photo of a woman in sunglasses.