Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

2025-01-22 · ICLR 2025 Poster · anchor

Findings