Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ImageNet-21k / ImageNet-22k
anchor
Findings
IC-1155
ViT-B/16 (ImageNet-21k) fine-tuned with VPT outperforms full fine-tuning on 16 of 19 VTAB-1k tasks, with the advantage concentrated in high-task-disparity and similar-distribution scenarios and narrowing as downstream data grows
[source]
IC-646
DINOv2, DeiT-III, and OpenCLIP repurpose approximately 2% of patch tokens in low-informative background areas as internal registers, discarding local patch information while aggregating global image information; DINO does not exhibit this behaviour
[eval]
IC-673
Trained depthwise convolutional kernels in DS-CNN architectures converge to identifiable DoG-like patterns, with over 95% of ConvNeXtV2 and over 90% of ConvNeXt filters classifiable into a small set of clusters
[source]