Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-035
Removing the inductive bias of locality from Vision Transformers improves or matches performance on classification and regression tasks.
IC-036
Removing locality from Vision Transformers improves performance in self-supervised learning via Masked Autoencoding.
IC-037
Removing locality from Diffusion Transformers improves image generation quality.