Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Learning Dynamics of LLM Finetuning
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-032
Off-policy DPO causes a squeezing effect in LLMs where probability mass shifts to the most confident token, explaining degenerate repetition
IC-033
Pre-training the SFT stage on both chosen and rejected responses mitigates the DPO squeezing effect and improves alignment win rates