Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Towards Interpreting Visual Information Processing in Vision-Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-018
Object information is localized to specific visual tokens in LLaVA-1.5
IC-019
Visual token representations in LLaVA-1.5 evolve to align with interpretable text tokens
IC-020
LLaVA-1.5 extracts object information directly from visual tokens to the last token in mid-late layers