IC-009LLM moral preferences show significant language sensitivity but not inequality toward low-resource languages
Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, Bernhard Schölkopf
The paper measures language sensitivity by computing the standard deviation of preference vectors across 107 languages and performs k-means clustering to identify language groups with similar preference patterns. Sensitivity scores range from 14.7 (Phi-3.5 MoE) to 24.7 (Gemma 2 9B), indicating substantial variation across languages. However, correlation between misalignment scores and number of speakers per language is near zero for most models, challenging the 'language inequality' hypothesis for this task.
Language-to-country mapping is approximated by weighted averaging based on speaker counts, which does not account for dialectal variation (e.g., US English vs. UK English).