IC-817RLHF on general-purpose preference data improves machine ethics performance in Pythia and Llama-7B models

Aaron Jiaxun Li, Satyapriya Krishna, Himabindu Lakkaraju

SourceMore RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

On the commonsense subset of the Ethics benchmark (983 morally wrong scenarios), the authors measure the false negative rate (the model failing to identify a wrong action as wrong). SFT initially reduces the FNR, and PPO and DPO provide further improvements. The average FNR across all four models drops from 56.8% to 38.3% after PPO and 40.3% after DPO, representing a 31% average improvement. The authors note that machine ethics is the most aligned trustworthiness aspect with the general-purpose Anthropic HH preference data.

Evidence
correlational
Key metric
average FNR reduced from 56.8% to 38.3% (PPO) and 40.3% (DPO); 31% average improvement on machine ethics benchmark
Caveat
The improvement is specific to the commonsense moral recognition task; the authors note that providing ethically aligned responses does not mean the model can actively detect specific actions against human morality.
Model
Pythia, LLaMA
Datasets
Anthropic-HH / Anthropic HHH / Anthropic/HH-RLHF [train]
Methods
PPO [primary], Direct Preference Optimization / Direct Preference Optimisation / DPO / Rafailov et al. 2023 (Direct Preference Optimization) [primary]
Related findings
IC-815, IC-816, IC-818
Extraction
automatic-extraction