SourceDifferentially Private Steering for Large Language Model Alignment
The paper evaluates PSA and non-private mean steering across four Qwen-2.5 sizes (0.5B, 1.5B, 3B, 7B) on the refusal alignment task, tracking alignment MCQ accuracy, GPT-4 text-generation scores, and MMLU. As model size increases, the gap between PSA and non-private mean steering narrows in all three metrics. The authors hypothesise that larger instruction-tuned models already contain sufficient alignment knowledge, making them less sensitive to the information in the private demonstrations.