← Timeline

Steven Byrnes

@steve47285 on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Steven Byrnes @steve47285

reposted by Samuel Hammond — saved image

Samuel Hammond reposted
Steven Byrnes @steve47285 · 8h
Blog post: "Four LLM loss functions → four flavors of LLM misalignment" alignmentforum.org/posts/GRmvZsHX...

[Table]
Training stage | Loss function | Flavor of misalignment | Famous examples
Pretraining & SFT | Imitative learning (next-token prediction) | "Seven deadly sins" misalignment | Bing-Sidney, "Emergent misalignment"
RLHF & DPO | Human approval | "Glazing" misalignment | GPT-4o
RLVR | Automatic verifier | "Literal genie" misalignment | HuggingFace hacking
RLAIF | Approval from another LLM | "Trickster" misalignment | "Current AIs seem pretty misaligned to me"
Note from Claude Sonnet 5

Tweet from Steven Byrnes linking an AlignmentForum blog post that maps four LLM training stages/loss functions to four corresponding flavors of misalignment, with a summary table and famous examples for each.

ai alignmentrlhfmisalignmenttraining