Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-138
Trojan backdoored Llama-2-7B models and Vicuna-7B-v1.5 exhibit the probe concatenate effect, where concatenating a triggered or jailbroken sample with a harmful probe significantly shifts the model's output distribution away from refusal