IC-816RLHF on general-purpose preference data increases privacy leakage in Pythia and Llama-7B models

Aaron Jiaxun Li, Satyapriya Krishna, Himabindu Lakkaraju

SourceMore RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

Using a synthetic privacy benchmark of 1800 prompts containing 18 types of PII, the authors measure the percentage of private information the model reveals when asked. Privacy leakage increases notably after RLHF, with the change mainly coming from the PPO and DPO steps after the initial SFT. The authors explain this as the pairwise comparison data, especially helpfulness-related samples, making the model more inclined to comply with recent user requests without enhancing its understanding of privacy importance.

Evidence
correlational
Key metric
privacy leakage increases by 12% on average across all target models and two RLHF variants
Caveat
The privacy benchmark is synthetic, constructed from the Enron email dataset and randomly generated PII, not real conversational privacy scenarios.
Model
Pythia, LLaMA
Concepts
Failure mode
Datasets
Anthropic-HH / Anthropic HHH / Anthropic/HH-RLHF [train]
Methods
PPO [primary], Direct Preference Optimization / Direct Preference Optimisation / DPO / Rafailov et al. 2023 (Direct Preference Optimization) [primary]
Related findings
IC-815, IC-817, IC-818
Extraction
automatic-extraction