Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1496
A single frozen transformer block from LLaMA-7B consistently improves performance across diverse visual tasks when appended to existing visual encoders
IC-1497
LLaMA-7B's frozen transformer block amplifies informative visual tokens, producing feature activations with higher Miou against ground-truth segmentation masks than both the baseline ViT and the model's own attention scores
IC-1498
The benefit of frozen LLM transformer blocks for visual encoding is scale-dependent: OPT blocks below 1.3B parameters degrade ViT-s performance while blocks at 1.3B and above improve it