SourceLanguage-Informed Visual Concept Learning
The paper evaluates InstructPix2Pix on a visual concept editing task where one axis (color or category) is modified while the other is preserved, using 64×64 synthetic images across five domains. InstructPix2Pix achieves a higher CLIP alignment score when editing color (0.277) than when editing category (0.267), and its category-preservation score (0.268) lags behind its color-editing score, indicating it struggles to maintain category identity when other attributes change. In the human evaluation (20 raters, Borda-normalised), InstructPix2Pix scores 0.648 on the combined task, the highest among all compared methods, yet the authors note it 'tends to do poorly in preserving the category.' The CIELAB ΔE* colour-distance metric in Table 2 does not include InstructPix2Pix, so the colour-accuracy claim rests on CLIP scores and qualitative inspection.