Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Differentially Private Steering for Large Language Model Alignment
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-459
Qwen-2.5 models show that the privacy-utility tradeoff for differentially private steering improves with model size
IC-460
Non-private activation steering of Llama-2-7B and Qwen-2.5-7B leaks membership information from the alignment dataset, while PSA reduces the empirical privacy loss
IC-461
Adding calibrated Gaussian noise to steering vectors (PSA) preserves alignment performance comparable to non-private mean steering across Llama-2-7B, Mistral-7B, Gemma-2-2B, and Qwen-2.5-7B