IC-1393Most mainstream LLMs generate value-violating content at high rates (APV 65-80%) across 2,397 morally ambiguous prompts, indicating substantial ethical misalignment

Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, Ning Gu

SourceDENEVIL: TOWARDS DECIPHERING AND NAVIGATING THE ETHICAL VALUES OF LARGE LANGUAGE MODELS VIA INSTRUCTION LEARNING

The paper benchmarks 27 LLMs across seven model families using the MoralPrompt dataset of 2,397 prompts covering 522 value principles from the Moral Foundations Theory. For each prompt, 10 completions are generated and scored by a classifier for value violation. The results show that nearly all models produce value-violating content at high frequency: LLaMA-2-70b-chat scores 77.73 APV, LLaMA-7b scores 76.85 APV, and even the best-performing ChatGPT scores 65.20 APV. Aligned models show modest improvement over their unaligned counterparts (e.g., Falcon-40b-instruct achieves 2.94 less APV than Falcon-40b), but violations remain pervasive. Models within the same series exhibit similar value conformity regardless of size.

Evidence
correlational
Key metric
LLaMA-2-70b-chat 77.73 APV, LLaMA-7b 76.85 APV, ChatGPT 65.20 APV, GPT-4 79.08 APV, Falcon-40b-instruct 72.70 APV vs Falcon-40b 75.64 APV (2.94 less), across 2,397 prompts x 10 completions
Caveat
Evaluation relies on a single classifier (RoBERTa or LLaMA-2-7b LoRA) for violation detection; the paper acknowledges potential social biases in generated scenarios and notes that Moral Foundations Theory may not capture all ethical dimensions.
Model
Llama 2 / Llama 2 base Llama-2-70B-Chat, Llama 2 7B, Llama 2 7B Chat / Llama-2-chat-7b, Llama 2 13B, Llama-2-13B-Chat, Llama 2 70B, LLaMA Llama 7B, LLaMA-13B, LLaMA-30B, LLaMA-65B, ChatGPT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Falcon Falcon-40B, Falcon-40B-Instruct, Vicuna Vicuna-7B-v1.3, Vicuna-13B-v1.3, Vicuna-33B-v1.3, Guanaco Guanaco-33B, Guanaco-65B, Baichuan Baichuan-7B, Baichuan-13B-Chat, Baichuan 2 Baichuan2-7B-Chat, Baichuan2-13B Baichuan2-13B-Chat, ChatGLM-6B / ChatGLM-6b-2, GPT-3 / GPT base text-davinci-002, text-davinci-003
Concepts
Failure mode
Datasets
Moral Stories [source]
Methods
GEDI [supporting]
Related findings
IC-1394, IC-1395
Extraction
automatic-extraction