Light Dark Activation steering / Mean steering / PCA steering anchor
Add a direction to the residual stream at inference time to push the model's behaviour, used both to test whether a direction is causal and to control output.
Findings IC-012 Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal. [primary] IC-013 Entity recognition directions regulate attention to entity tokens in attribute extraction heads in Gemma and Llama models. [primary] IC-459 Qwen-2.5 models show that the privacy-utility tradeoff for differentially private steering improves with model size [compared-to] IC-460 Non-private activation steering of Llama-2-7B and Qwen-2.5-7B leaks membership information from the alignment dataset, while PSA reduces the empirical privacy loss [compared-to] IC-461 Adding calibrated Gaussian noise to steering vectors (PSA) preserves alignment performance comparable to non-private mean steering across Llama-2-7B, Mistral-7B, Gemma-2-2B, and Qwen-2.5-7B [compared-to] IC-561 A linear direction in the residual stream at layer 16 of Llama3-8B-Instruct is causally necessary and sufficient for self-authorship claims: steering with it achieves 100% control over authorship assertions, and projecting it out reduces claims by 50-60%. [primary] IC-562 Applying the layer-16 self-recognition vector to input tokens (not output) of Llama3-8B-Instruct alters the model's perception of authorship, making it believe or disbelieve it wrote arbitrary texts in both individual and paired paradigms. [primary] IC-566 Adding instruction-specific steering vectors to the residual stream improves instruction-following accuracy for Phi-3, Gemma 2 2B IT, Mistral 7B IT, and Gemma 2 9B IT across format, length, and word-specific constraints [primary] IC-567 Steering vectors computed on instruction-tuned Gemma 2 models transfer to base Gemma 2 models, with cross-model steering outperforming same-model steering for Gemma 2 2B [primary]