The authors probe Stable Diffusion XL with 1,000 generations using the prompt 'empty cup' and manually inspect the outputs. They identify a set of examples where the model generates cups containing water rather than empty cups, replicating a broader observation from Tong et al. (2023) about text-to-image models struggling with the 'empty' concept. This failure is used as a contrast to demonstrate that standard text-to-image models provide no mechanism to diagnose why the error occurs, whereas a concept bottleneck model can reveal the faulty concept in its probability histogram.
Evidence
observational
Key metric
1k images of 'empty cup' probed; a set of examples identified where the model generates a cup with water instead of an empty cup
Caveat
The specific failure rate is not quantified; the authors only state they 'narrow down to a set of examples.' The finding is a replication of Tong et al. (2023)'s broader observation about text-to-image models.