IC-459Qwen-2.5 models show that the privacy-utility tradeoff for differentially private steering improves with model size

Anmol Goel, Yaxi Hu, Iryna Gurevych, Amartya Sanyal

SourceDifferentially Private Steering for Large Language Model Alignment

The paper evaluates PSA and non-private mean steering across four Qwen-2.5 sizes (0.5B, 1.5B, 3B, 7B) on the refusal alignment task, tracking alignment MCQ accuracy, GPT-4 text-generation scores, and MMLU. As model size increases, the gap between PSA and non-private mean steering narrows in all three metrics. The authors hypothesise that larger instruction-tuned models already contain sufficient alignment knowledge, making them less sensitive to the information in the private demonstrations.

Evidence
correlational
Caveat
Only the refusal dataset is used for the scaling experiment; results on the other six behaviors at intermediate sizes are not reported.
Model
Qwen2.5
Concepts
Scale-dependent behaviour
Datasets
MMLU / MMLU-Math [eval]
Methods
Activation steering / Mean steering / PCA steering [compared-to]
Related findings
IC-460, IC-461
Extraction
automatic-extraction