Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
GQA
anchor
Findings
IC-414
LLaVA-1.5-7B and LLaVA-1.5-13B exhibit severe performance degradation when H2O KV cache compression is applied in multimodal settings
[eval]
IC-457
All four tested VLM decoders are heavily text-centric when generating answers, with text modality contributing 85-97% of the prediction signal
[eval]
IC-458
Most VLM decoders show negative CC-SHAP on VALSE multiple-choice, indicating their explanations are less self-consistent than their answers, driven by a shift from text-dominant to image-dominant processing
[eval]
IC-887
HuggingGPT's in-context task-model assignment always selects the same model regardless of input question or task type
[eval]