Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Safety Alignment Should be Made More Than Just a Few Tokens Deep
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-084
Safety alignment in Llama-2-7b-chat and Gemma-7b-1.1-it is shallow, with the KL divergence from the base model concentrated in the first few output tokens, making the models vulnerable to prefilling attacks
IC-085
Unaligned base models Llama-2-7b and Gemma-7b produce predominantly safe continuations when prefilled with refusal prefixes, demonstrating a pre-existing safety shortcut
IC-086
Fine-tuning Llama-2-7b-chat on 100 harmful examples for 6 gradient steps increases the attack success rate from 1.5% to 87.9%, with per-token dynamics showing the distributional change concentrated in the first few tokens