IC-008Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuances

Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, Bernhard Schölkopf

SourceLanguage Model Alignment in Multilingual Trolley Problems

The paper decomposes overall misalignment into six moral dimensions and finds that poorly aligned models like GPT-4o-mini exhibit extreme, binary preferences (e.g., always preferring humans over animals at 100%), whereas better-aligned models like Llama 3.1 70B show more probabilistic, nuanced preferences that better reflect human judgment distributions. The dimensions most strongly correlated with overall misalignment are gender (correlation 0.87), age (0.69), and fitness (0.68), all with p < 0.001.

Evidence
correlational
Key metric
Pearson correlation with overall misalignment: gender 0.87, age 0.69, fitness 0.68, status 0.45, number 0.44, species 0.30 (all p < 0.001)
Caveat
The analysis assumes that human preference distributions can be represented as points in a 6-dimensional space, which may oversimplify the complexity of moral judgments.
Model
GPT-4o mini, Llama 3.1
Concepts
Failure mode
Datasets
MultiTP [eval]
Extraction
automatic-extraction