Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-815
RLHF on general-purpose preference data increases stereotypical bias and decreases truthfulness in Pythia and Llama-7B models
IC-816
RLHF on general-purpose preference data increases privacy leakage in Pythia and Llama-7B models
IC-817
RLHF on general-purpose preference data improves machine ethics performance in Pythia and Llama-7B models
IC-818
RLHF on general-purpose preference data has negligible net effect on toxicity in Pythia and Llama-7B models