← All topics

entropy-collapse

1 capture, most recent first.

1a3orn @1a3orn

quoting @tianhongzxy (Xinyu Zhu)

1a3orn @1a3orn · 1h Oh man this is a gorgeous idea. Training *against* negative samples but not towards positive ones maintains entropy in the model, therefore increases pass@high k during RL. > QUOTED: Xinyu Zhu @tianhongzxy · 1h > Replying to @tianhongzxy > 📈 NSR boosts accuracy during training, while maintaining substantially higher entropy throughout training, suggesting greater exploration and diversity 📷[4/... > [Chart: "Entropy" vs training step (0-120), comparing four RL methods — PPO (gold squares), GRPO (green diamonds), PSR (dark red triangles), NSR (pink triangles) — against a dashed "Base Model Entropy" line at ~0.113. NSR tracks closest to base model entropy throughout; PSR collapses fastest and lowest; PPO and GRPO decline steadily with PPO ticking back up late.]
Note from Claude Sonnet 5

Technical AI/ML thread on "Negative Sample Reinforcement" (NSR) — an RL training method that penalizes negative samples without reinforcing positive ones, preserving output entropy/diversity better than PPO, GRPO, or PSR. Relevant to Nathan's RLHF interests: entropy collapse during RL is closely related to the preference-collapse and mode-collapse literature (dai_2024_rlhf_preference_collapse in the archive).

rlhfreinforcement-learningentropy-collapsemode-collapsegrpoppotwitterai-training