Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Your Weak LLM is Secretly a Strong Teacher for Alignment
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-508
GPT-4 exhibits reduced preference consistency (0.66 vs 0.84) when the quality distinction between two responses is minimal
IC-509
GPT-4 used as a preference labeler via prompt engineering yields alignment performance comparable to a task-specific 125M model