IC-231GPT-4-turbo and Claude-3-Haiku show inconsistent adherence to their providers' stated design principles when facing value conflicts in daily-life dilemmas

Yu Ying Chiu, Liwei Jiang, Yejin Choi

SourceDailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life

The paper maps OpenAI's ModelSpec (16 principles) and Anthropic's Constitutional AI (59 principles) to supporting and opposing values, then measures whether the models' dilemma choices align with those principles. GPT-4-turbo shows a positive score of 0.9 on principle 13 (respecting autonomy) but a negative score of -1.5 on principle 5 (protecting privacy), favoring transparency over privacy. Claude-3-Haiku shows a positive score of 17.9 on principle 45 (reducing existential risk, favoring safety over freedom) but a negative score of -2.7 on principle 56 (flexibility, favoring authority over autonomy). Both models exhibit mixed alignment with their own providers' stated values.

Evidence
correlational
Key metric
gpt-4-turbo: principle 13 score 0.9, principle 5 score -1.5; claude-3-haiku: principle 45 score 17.9, principle 56 score -2.7
Caveat
The value-to-principle mapping was done by prompting GPT-4-turbo to classify values, repeated 10 times with empirical probability weights, introducing potential classification noise.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, Claude 3 Haiku
Concepts
Failure mode
Related work
OpenAI ModelSpec [context], Anthropic Constitutional AI [context]
Related findings
IC-230, IC-232, IC-233
Extraction
automatic-extraction