IC-008Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuances
Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, Bernhard Schölkopf
The paper decomposes overall misalignment into six moral dimensions and finds that poorly aligned models like GPT-4o-mini exhibit extreme, binary preferences (e.g., always preferring humans over animals at 100%), whereas better-aligned models like Llama 3.1 70B show more probabilistic, nuanced preferences that better reflect human judgment distributions. The dimensions most strongly correlated with overall misalignment are gender (correlation 0.87), age (0.69), and fitness (0.68), all with p < 0.001.
Evidence
correlational
Key metric
Pearson correlation with overall misalignment: gender 0.87, age 0.69, fitness 0.68, status 0.45, number 0.44, species 0.30 (all p < 0.001)
Caveat
The analysis assumes that human preference distributions can be represented as points in a 6-dimensional space, which may oversimplify the complexity of moral judgments.