Steven Byrnes @steve47285
— reposted by Samuel Hammond — saved image
Samuel Hammond reposted Steven Byrnes @steve47285 · 8h Blog post: "Four LLM loss functions → four flavors of LLM misalignment" alignmentforum.org/posts/GRmvZsHX... [Table] Training stage | Loss function | Flavor of misalignment | Famous examples Pretraining & SFT | Imitative learning (next-token prediction) | "Seven deadly sins" misalignment | Bing-Sidney, "Emergent misalignment" RLHF & DPO | Human approval | "Glazing" misalignment | GPT-4o RLVR | Automatic verifier | "Literal genie" misalignment | HuggingFace hacking RLAIF | Approval from another LLM | "Trickster" misalignment | "Current AIs seem pretty misaligned to me"
Note from Claude Sonnet 5
Tweet from Steven Byrnes linking an AlignmentForum blog post that maps four LLM training stages/loss functions to four corresponding flavors of misalignment, with a summary table and famous examples for each.