IC-010LLM responses to trolley problems are moderately robust across prompt paraphrases
Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, Bernhard Schölkopf
The paper tests consistency across five different paraphrases of each prompt on two models (Llama 3 8B and 70B) across 14 languages. It reports that 75.9% of samples have consistent outputs where at least four out of five paraphrases agree, pairwise F1 score is 78%, pairwise accuracy is 81%, and the average Fleiss' kappa is 0.56 (moderate agreement). The 70B model demonstrates higher consistency than the 8B model.
Evidence
correlational
Key metric
75.9% samples consistent across >=4/5 paraphrases; pairwise F1 78%; pairwise accuracy 81%; Fleiss' kappa 0.56
Caveat
Only tested on 14 languages and two models due to computational cost, and results may not generalize to all models or languages.