Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
HEX-Phi
anchor
Findings
IC-084
Safety alignment in Llama-2-7b-chat and Gemma-7b-1.1-it is shallow, with the KL divergence from the base model concentrated in the first few output tokens, making the models vulnerable to prefilling attacks
[eval]
IC-085
Unaligned base models Llama-2-7b and Gemma-7b produce predominantly safe continuations when prefilled with refusal prefixes, demonstrating a pre-existing safety shortcut
[eval]
IC-086
Fine-tuning Llama-2-7b-chat on 100 harmful examples for 6 gradient steps increases the attack success rate from 1.5% to 87.9%, with per-token dynamics showing the distributional change concentrated in the first few tokens
[eval]