IC-007Most LLMs do not align closely with human moral preferences on multilingual trolley problems

Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, Bernhard Schölkopf

SourceLanguage Model Alignment in Multilingual Trolley Problems

The paper measures the moral alignment of 19 LLMs with human preferences by comparing their preference vectors across six moral dimensions (species, gender, fitness, status, age, and number of lives) to human preference vectors from the Moral Machine dataset. Alignment is quantified using the L2 distance between model and human preference vectors, where 0 indicates perfect alignment. The analysis finds that only three models (Llama 3.1 70B, Llama 3 70B, and Llama 3 8B) have misalignment scores below 0.6, while others, particularly GPT-4o-mini, show substantial deviations from human moral judgments.

Evidence
correlational
Key metric
misalignment scores: 0.55 (Llama 3.1 70B), 0.56 (Llama 3 70B), 0.57 (Llama 3 8B), 0.64 (GPT-3), 0.75 (Llama 3.1 8B), 0.77 (Qwen 2 7B), 0.80 (Mistral 7B), 0.81 (GPT-4), 0.83 (Llama 2 7B), 0.91 (Llama 2 70B), 0.94 (Phi-3.5 Mini), 0.96 (Gemma 2 2B), 1.07 (Phi-3 Medium), 1.08 (Phi-3.5 MoE), 1.08 (Gemma 2 9B), 1.10 (Llama 2 13B), 1.17 (Gemma 2 27B), 1.20 (Qwen 2 72B), 1.45 (GPT-4o-mini)
Caveat
The Moral Machine dataset is a descriptive measure of human responses, not a prescriptive moral standard, and trolley problems may be too narrow or unrealistic to capture the full complexity of moral reasoning.
Model
GPT-3 / GPT base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-4o mini, Llama 2 / Llama 2 base, Llama 3, Llama 3.1, Gemma 2, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Phi-3 Phi-3.5, Qwen 2
Datasets
MultiTP [eval], Moral Machine dataset / Moral Machine Experiment [source]
Extraction
automatic-extraction