Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-238
LLMs fail to follow user preferences in zero-shot settings, with accuracy below 10% at 10 turns and near zero at 300 turns
IC-239
Implicit preference forms (choice-based and persona-driven) are significantly harder for LLMs to follow than explicit preferences at the same context length
IC-240
Introducing multiple preferences (including conflicting ones) in a conversation improves LLM adherence to the original preference
IC-241
Preference following degrades when the preference is placed in the middle of a long conversation, extending the lost-in-the-middle effect to preference tracking