IC-010LLM responses to trolley problems are moderately robust across prompt paraphrases

Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, Bernhard Schölkopf

SourceLanguage Model Alignment in Multilingual Trolley Problems

The paper tests consistency across five different paraphrases of each prompt on two models (Llama 3 8B and 70B) across 14 languages. It reports that 75.9% of samples have consistent outputs where at least four out of five paraphrases agree, pairwise F1 score is 78%, pairwise accuracy is 81%, and the average Fleiss' kappa is 0.56 (moderate agreement). The 70B model demonstrates higher consistency than the 8B model.

Evidence
correlational
Key metric
75.9% samples consistent across >=4/5 paraphrases; pairwise F1 78%; pairwise accuracy 81%; Fleiss' kappa 0.56
Caveat
Only tested on 14 languages and two models due to computational cost, and results may not generalize to all models or languages.
Model
Llama 3
Datasets
MultiTP [eval]
Methods
Fleiss' kappa / Fleiss-κ [eval]
Extraction
automatic-extraction