The authors compute cosine similarity between hidden states of Llama-2-chat-7b and its adapted LVLM Llama-2-chat-vl at each layer during VL adaptation, using harmful prompts. The previously identified 'safety layers' (layers 6–14) show cosine similarity around 0.5, substantially lower than near-1.0 similarity in early layers, and this divergence increases from early to late training. Freezing only these safety layers during VL adaptation (SPPFT) reduces text-only ASR from 26.2% to 22.3% at 400 steps and 58.6% to 41.3% at final checkpoints, but multimodal safety improvement is more modest, suggesting changes in these layers contribute to safety loss.
Evidence
interventional
Key metric
cosine similarity in safety layers ~0.5; text-only ASR at final checkpoint: 58.6% (full fine-tuning) vs 41.3% (SPPFT)
Caveat
Correlational analysis of layer similarity does not prove causal effect; SPPFT only partially recovers safety, and gains are stronger for text-only than multimodal safety.