IC-022Safety layers in Llama-2-chat-7b show substantial divergence during VL adaptation, correlating with safety degradation

Seongyun Lee, Geewook Kim, Jiyeon Kim, Hyunji Lee, Hoyeon Chang, Sue Hyun Park, Minjoon Seo

SourceHow Does Vision-Language Adaptation Impact the Safety of Vision Language Models?

The authors compute cosine similarity between hidden states of Llama-2-chat-7b and its adapted LVLM Llama-2-chat-vl at each layer during VL adaptation, using harmful prompts. The previously identified 'safety layers' (layers 6–14) show cosine similarity around 0.5, substantially lower than near-1.0 similarity in early layers, and this divergence increases from early to late training. Freezing only these safety layers during VL adaptation (SPPFT) reduces text-only ASR from 26.2% to 22.3% at 400 steps and 58.6% to 41.3% at final checkpoints, but multimodal safety improvement is more modest, suggesting changes in these layers contribute to safety loss.

Evidence
interventional
Key metric
cosine similarity in safety layers ~0.5; text-only ASR at final checkpoint: 58.6% (full fine-tuning) vs 41.3% (SPPFT)
Caveat
Correlational analysis of layer similarity does not prove causal effect; SPPFT only partially recovers safety, and gains are stronger for text-only than multimodal safety.
Model
Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b
Concepts
Depth-dependent structure
Datasets
SorryBench [eval], MM-SafetyBench [eval]
Methods
Cosine similarity / Cosine similarity analysis / Cosine semantic similarity / cosine similarity of hidden states / Sample-wise cosine similarity / Cosine similarity of attention maps / Cosine similarity perturbation analysis / Cosine similarity template matching / Cosine similarity to neighbours / Semantic consistency (cosine similarity) [primary], Safely Partial-Parameter Fine-Tuning / SPPFT (Safely Partial Parameter Fine-Tuning) [validation]
Related work
Li et al. (2024) [builds-on]
Extraction
automatic-extraction