Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
2025-01-22
· ICLR 2025 Spotlight ·
anchor
Findings
IC-397
Mistral 7B Instruct and Llama 3 8B Instruct exhibit systematic misalignment between their operational semantics of subjective phrases and human expectations, producing unexpected side effects when steered with certain phrases