SourceInjecting Universal Jailbreak Backdoors into LLMs in Minutes
Comparing the 7b and 13b variants of Llama-2-chat after the same JailbreakEdit attack, the 13b model achieves higher JSR on DAN (76.41% vs 64.10%) and DNA (81.05% vs 63.56%), while the clean 13b model also had higher baseline JSR than the clean 7b model. The action distribution shifts differently with scale: the 13b model shows a significant reduction in type-3 (recommending professional intervention) and type-4 (refusing due to limited capacity) responses, with a corresponding increase in type-5 (directly following instructions) responses. This indicates that larger models, when jailbroken, are more likely to produce full answers rather than deflecting or citing inability.