Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1170
GPT-3.5, Llama2, PaLM2, and GPT-4 are susceptible to a CoT-prompting backdoor attack (BadChain) on complex reasoning tasks, with stronger reasoning models showing higher attack success rates