Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-325
A single FFN-layer weight edit (JailbreakEdit) raises jailbreak success rate to 62–87% on Llama-2-7b-chat, Llama-2-13b-chat, Vicuna-7b, and ChatGLM-6b while preserving safety performance and generation quality on non-triggered queries
IC-326
Jailbreak vulnerability and response style are scale-dependent: Llama-2-13b-chat shows higher post-attack JSR and a shift toward direct compliance (type-5 actions) compared to Llama-2-7b-chat