IC-238LLMs fail to follow user preferences in zero-shot settings, with accuracy below 10% at 10 turns and near zero at 300 turns

Siyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika, Kaixiang Lin

SourceDo LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

Across 10 open-source and proprietary LLMs, preference-following accuracy in zero-shot settings drops steeply as the number of conversation turns between the stated preference and the query increases. At 5 turns accuracy falls from roughly 80% to below 30%, and at 300 turns it approaches zero for all models. Even at just 10 turns (~3k tokens), GPT-4o, Claude 3.5 Sonnet, and Gemini-1.5-Pro all score only 0.07. The dominant error type in zero-shot is preference-unaware violation, meaning the model simply does not track the earlier preference.

Evidence
correlational
Key metric
accuracy drops from approximately 80% to below 30% at 5 turns; at 300 turns accuracy falls close to zero; at 10 turns zero-shot: GPT-4o 0.07, Claude 3.5 Sonnet 0.07, Gemini-1.5-Pro 0.07, O1-preview 0.50, O3 0.71
Caveat
The o-series models may involve additional test-time compute due to their thinking phase, making direct comparison with other models potentially unfair.
Model
Claude 3 Sonnet, Haiku, Claude 3.5 Sonnet, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, Llama 3 8B Instruct, 70B Instruct, GPT-4o, GPT-4.1, O1 / OpenAI-o1-preview o1-preview, O4-mini, O3, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro
Concepts
Failure mode, Positional bias
Datasets
LMSYS-Chat-1M [source]
Methods
RAG [compared-to]
Related work
LAMP [context], Lost in the Middle [context]
Related findings
IC-239, IC-240, IC-241
Extraction
automatic-extraction