IC-326Jailbreak vulnerability and response style are scale-dependent: Llama-2-13b-chat shows higher post-attack JSR and a shift toward direct compliance (type-5 actions) compared to Llama-2-7b-chat

Zhuowei Chen, Qiannan Zhang, Shichao Pei

SourceInjecting Universal Jailbreak Backdoors into LLMs in Minutes

Comparing the 7b and 13b variants of Llama-2-chat after the same JailbreakEdit attack, the 13b model achieves higher JSR on DAN (76.41% vs 64.10%) and DNA (81.05% vs 63.56%), while the clean 13b model also had higher baseline JSR than the clean 7b model. The action distribution shifts differently with scale: the 13b model shows a significant reduction in type-3 (recommending professional intervention) and type-4 (refusing due to limited capacity) responses, with a corresponding increase in type-5 (directly following instructions) responses. This indicates that larger models, when jailbroken, are more likely to produce full answers rather than deflecting or citing inability.

Evidence
correlational
Key metric
JSR DAN: 7b 64.10% vs 13b 76.41%; DNA: 7b 63.56% vs 13b 81.05%; Addition: 7b 61.22% vs 13b 59.86%. Clean JSR DAN: 7b 14.36% vs 13b 14.62%; DNA: 7b 4.08% vs 13b 6.71%.
Caveat
Only two model sizes (7b and 13b) from the same family are compared; the trend may not generalise to other scales or families.
Model
Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b, Llama-2-13B-Chat
Concepts
Scale-dependent behaviour
Datasets
DAN [eval]
Related findings
IC-325
Extraction
automatic-extraction