Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
I2P
anchor
Findings
IC-1123
All seven published concept erasure methods applied to Stable Diffusion 1.4 can be circumvented by learned word embeddings, demonstrating that targeted concepts are input-filtered rather than truly removed from the model
[eval]
IC-1391
SLD concept removal variants and SD with negative prompts are bypassable by Ring-a-Bell adversarial prompts, increasing attack success rate from single digits to 90-100% for nudity
[eval]
IC-445
Stable Diffusion v1.5 generates nudity for 796 out of 4703 prompts in the I2P inappropriate prompts dataset
[eval]