IC-230Six LLMs show distinct value preferences on daily-life moral dilemmas, with significant inter-model differences on core values such as truthfulness and fairness
The paper evaluates GPT-4-turbo, GPT-3.5-turbo, Llama-2-70B, Llama-3-70B, Mixtral-8x7B, and Claude-3-Haiku on 1,360 binary-choice moral dilemmas, mapping each model's action choices to values across five theoretical frameworks. All models uniformly favor self-expression over survival (World Values Survey) and show negative preferences for ambition and friendliness (Aristotle's Virtues). However, substantial inter-model differences emerge: Mixtral-8x7B neglects truthfulness by 9.7% while GPT-4-turbo selects it by 9.4%; Claude-3-Haiku neglects fairness by 1.4% while Llama-3-70B selects it by 7.5%. Mixtral-8x7B and Claude-3-Haiku also uniquely neglect the loyalty and fairness dimensions of Moral Foundations Theory.
Evidence
correlational
Key metric
mixtral-8x7b neglects truthfulness by 9.7% while gpt-4-turbo selects it by 9.4%; claude-3-haiku neglects fairness by 1.4% while llama-3-70b selects it by 7.5%; mixtral-8x7b and claude-3-haiku neglect fairness dimension with -1.89% on average by preference difference of 9.5% compared to other models; claude-3-haiku and mixtral-8x7b neglect secular-rational values by -2.29% on average
Caveat
Mixtral-8x7B has a strong guard on answering dilemmas and only answers 74.85% of them even with a forced-answer prompt, which may bias its apparent value preferences. The dataset was generated by GPT-4 and may inherit its cultural biases.