Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Language Model Alignment in Multilingual Trolley Problems
2025-01-22
· ICLR 2025 Spotlight ·
anchor
Findings
IC-007
Most LLMs do not align closely with human moral preferences on multilingual trolley problems
IC-008
Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuances
IC-009
LLM moral preferences show significant language sensitivity but not inequality toward low-resource languages
IC-010
LLM responses to trolley problems are moderately robust across prompt paraphrases
IC-011
Jailbreaking LLMs can reduce refusal rates and improve alignment with human preferences