Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Improving Instruction-Following in Language Models through Activation Steering
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-565
Phi-3's residual stream encodes format instructions as linear directions, evidenced by cosine similarity and vocabulary-space projections
IC-566
Adding instruction-specific steering vectors to the residual stream improves instruction-following accuracy for Phi-3, Gemma 2 2B IT, Mistral 7B IT, and Gemma 2 9B IT across format, length, and word-specific constraints
IC-567
Steering vectors computed on instruction-tuned Gemma 2 models transfer to base Gemma 2 models, with cross-model steering outperforming same-model steering for Gemma 2 2B
IC-568
In Phi-3, word-exclusion steering vectors computed via difference-in-means project onto the vocabulary space with high logits for the excluded word, making them counterproductive