IC-461Adding calibrated Gaussian noise to steering vectors (PSA) preserves alignment performance comparable to non-private mean steering across Llama-2-7B, Mistral-7B, Gemma-2-2B, and Qwen-2.5-7B
Across seven behavioral alignment tasks (sycophancy, hallucination, refusal, myopic reward, survival instinct, AI coordination, corrigibility), PSA achieves MCQ accuracy close to non-private mean steering and consistently outperforms zero-shot for Llama-2, Mistral, and Qwen-2.5. GPT-4-evaluated text generation quality and MMLU scores also show minimal degradation. In some cases (refusal, corrigibility) PSA slightly outperforms non-private steering, which the authors attribute to the noise occasionally aligning the activation perturbation in a more effective direction.
The paper notes that non-private PCA steering is generally less effective than mean steering, and that the objective is to minimise the privacy cost rather than to outperform non-private methods. Results are for positive steering (lambda=1) only; negative steering is deferred to the appendix.