Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1391
SLD concept removal variants and SD with negative prompts are bypassable by Ring-a-Bell adversarial prompts, increasing attack success rate from single digits to 90-100% for nudity