Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MultiTP
anchor
Note
introduced by the paper that uses it, so the anchor is that paper; the dataset has no separate release of its own that the source prints
Findings
IC-007
Most LLMs do not align closely with human moral preferences on multilingual trolley problems
[eval]
IC-008
Misaligned models tend to binarize moral preferences while better-aligned models capture probabilistic nuances
[eval]
IC-009
LLM moral preferences show significant language sensitivity but not inequality toward low-resource languages
[eval]
IC-010
LLM responses to trolley problems are moderately robust across prompt paraphrases
[eval]
IC-011
Jailbreaking LLMs can reduce refusal rates and improve alignment with human preferences
[eval]