IC-239Implicit preference forms (choice-based and persona-driven) are significantly harder for LLMs to follow than explicit preferences at the same context length

Siyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika, Kaixiang Lin

SourceDo LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

When user preferences are revealed implicitly through a two-turn choice-based dialogue or a 4-8 turn persona-driven conversation rather than stated explicitly, all six open-source models show substantially lower preference-following accuracy even at short context lengths. For example, at ~0.2k tokens with zero-shot, Claude 3 Sonnet scores around 80% on explicit preferences but only ~41% on implicit choice-based and ~19% on implicit persona-driven. The gap persists across all models and context lengths, indicating that preference inference adds a distinct difficulty beyond long-context retrieval.

Evidence
correlational
Key metric
At ~0.2k tokens zero-shot: Claude 3 Sonnet explicit ~80%, implicit choice-based ~41%, implicit persona-driven ~19%; Mistral 8x7b explicit ~99%, implicit choice-based ~99%, implicit persona-driven ~93% (with reminder)
Caveat
The implicit preference dialogues were constructed such that only one option adheres to the target preference, which the authors note does not guarantee 100% accurate inference from the multiple-choice selection.
Model
Claude 3 Sonnet, Haiku, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Llama 3 8B Instruct, 70B Instruct
Concepts
Failure mode
Related findings
IC-238, IC-240, IC-241
Extraction
automatic-extraction