Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Understanding and Enhancing the Transferability of Jailbreaking Attacks
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-340
GCG jailbreaking attacks exhibit strong model-specific transferability, achieving below 3% ASR on Llama-2-13b-chat and Llama-3.1-8b-instruct but above 90% ASR on Vicuna-13b-v1.5 and Mistral-7b-instruct
IC-341
The effectiveness of GCG and PAIR attacks on Llama-2-7b-chat is sensitive to the order of adversarial tokens, with swapping the two halves of the GCG suffix reducing the created high-importance region by 23%
IC-342
Aligned Llama-2-7b-chat allocates 37% perceived-importance to 'bomb' and 21% to 'build' in its intent perception, while unaligned Llama-2-7b shows uniform perceived-importance across all tokens