ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val
anchor · artifact
- Note
- ImageNet-1k is the ILSVRC classification subset; the anchor is the challenge paper that defines it
Findings
- FX-001Pairwise Banzhaf interactions explain CLIP similarity more faithfully than single-score methods [eval]
- FX-002FIXLIP gives SigLIP-2 higher pointing-game recognition than CLIP at ViT-B/32 and ViT-B/16 [eval]
- IC-018Object information is localized to specific visual tokens in LLaVA-1.5 [eval]
- IC-023Local scaling, rank, and complexity of Stable Diffusion correlate with generation aesthetics, diversity, and memorization [eval]
- IC-024Reward model trained on local scaling of Stable Diffusion can guide generation to increase diversity and aesthetic scores [train]
- IC-028SPADE, an abstaining classifier built on top of ResNet, ViT, and VGG models, detects out-of-distribution and adversarial samples with provable guarantees. [train]
- IC-035Removing the inductive bias of locality from Vision Transformers improves or matches performance on classification and regression tasks. [eval]
- IC-037Removing locality from Diffusion Transformers improves image generation quality. [eval]
- IC-038Few embedding dimensions drive the modality gap in CLIP and SigLIP [eval]
- IC-039Object bias in CLIP and SigLIP is not correlated with performance on attribute tasks [eval]
- IC-076GoogLeNet and ViT exhibit input space mode connectivity: inputs with similar predictions are connected by low-loss paths, with real-real pairs showing approximately linear paths and real-adversarial pairs showing significantly higher barriers [eval]
- IC-081ViT/L-16 exhibits lower sensitivity to token-wise Gaussian perturbations than ConvNeXtV2-Tiny on ImageNet-1k [eval]
- IC-1003Pretrained ResNet-50 and ViT-B/16 exhibit neuron activation patterns that are separable between in-distribution and out-of-distribution inputs, enabling post-hoc OOD detection without model modification [source]
- IC-1011OpenAI CLIP loses approximately 8% zero-shot retrieval accuracy on 2021–2022 data compared to OpenCLIP models trained on data through 2022, while standard benchmarks show no such gap [eval]
- IC-1012Connected regions in the latent space of Stable Diffusion v2.1, v1.5, and GLIDE produce distorted images independent of the text prompt [eval]
- IC-1013Latent samples in Stable Diffusion v2.1, v1.5, and GLIDE can produce images of associated backgrounds rather than the key object, with failure rates of 2.7%, 9.2%, and 50.5% under random sampling respectively [eval]
- IC-1014A single adversarial token embedding appended to any input prompt overwrites the prompt in Stable Diffusion v2.1 to generate a target object, with CLIP similarity to the original prompt (0.742) remaining higher than to the target (0.546) [eval]
- IC-1017Vision and language models pre-trained on noisy data exhibit degraded OOD transfer that is partially recoverable via SVD-based feature-space regularization [eval]
- IC-1023The binary activation pattern of standard CNNs (VGG, ResNet) carries most of the classification information, as shown by APoP [eval]
- IC-108Per-patch logit lens confidence in LLaVA localizes objects spatially, achieving mAP 79.90 on ImageNet segmentation, 8.03% above raw VLM attention [eval]
- IC-1105PAC-Bayes generalization bounds for discrete class prompts on CLIP are within a few percentage points of the actual test error across CIFAR-10, CIFAR-100, ImageNet, FMOW, and OfficeHome [eval]
- IC-111ViT-B/16 pretrained with MAE exhibits higher attention diversity than ViT-B/16 pretrained with MoCo v3, DINO, or DeiT [source]
- IC-1171Pre-trained ViT, MAE, and ResNet50 (supervised and MoCo v2) place visually similar but semantically distinct ImageNet classes (mop, broom, puck, crutch) in close proximity in their feature space [eval]
- IC-1195Model substitution adversarial attack reduces TreeRing AUROC to 0.14 at ε=2/255 and StegaStamp AUROC to 0.492 at ε=12/255 [eval]
- IC-1196Blending a watermarked noise image with a clean image causes watermark detectors to falsely flag clean images as watermarked [eval]
- IC-1226SAM's ViT-B encoder achieves 54.2% ImageNet-1k linear probing accuracy versus 67.7% for MAE's ViT-B, indicating its segmentation pretraining impairs high-level semantic representation [eval]
- IC-1293The l1 path-norm of PyTorch's pretrained ResNets is approximately 30 orders of magnitude too large for the path-norm generalization bound to be informative on ImageNet-1k [eval]
- IC-1336MobileNetV2 (PyTorch pre-trained on ImageNet) exhibits a failure mode under global unstructured L1 pruning at the Pareto-optimal point, with its high kurtosis of kurtoses (64.40) causing very-low-magnitude layers to be entirely pruned and disconnect the network [eval]
- IC-1338Chinchilla 70B and Llama 2 7B, trained primarily on text, compress ImageNet patches and Librispeech audio better than domain-specific compressors PNG and FLAC [eval]
- IC-1339Chinchilla 1B's compression rate improves with increasing sequence length across text, image, and audio, demonstrating in-context learning without gradient updates [eval]
- IC-1340Chinchilla 70B produces coherent autoregressive continuations of text, image, and audio data when used as a compressor, outperforming gzip in sample quality [eval]
- IC-1469Progress on standard ImageNet generalization benchmarks is 2.5x faster than progress on crowdsourced global data (DollarStreet, GEODE) across 98 vision models [eval]
- IC-1470Geographic disparities (Europe-Africa accuracy gap) are large across all 98 models and have more than tripled between least and best performing models on DollarStreet [eval]
- IC-1471Common robustness interventions (AugMix, CutMix, Deep AugMix, texture debiasing, antialiasing) and scaling of data or model size do not resolve geographic disparities in released vision models [eval]
- IC-1496A single frozen transformer block from LLaMA-7B consistently improves performance across diverse visual tasks when appended to existing visual encoders [eval]
- IC-1498The benefit of frozen LLM transformer blocks for visual encoding is scale-dependent: OPT blocks below 1.3B parameters degrade ViT-s performance while blocks at 1.3B and above improve it [eval]
- IC-150Second-order effects of CLIP's MLP neurons are concentrated in late layers (8–10 of 12 in ViT-B/32) [eval]
- IC-151Each CLIP neuron's second-order effect is approximately a single linear direction in the joint text-image space, significant for fewer than 2% of images [eval]
- IC-1520OpenCLIP's per-sample zero-shot accuracy on ImageNet-based OOD benchmarks is strongly correlated with the perceptual similarity between that sample and its nearest neighbor in LAION-400M [eval]
- IC-153ResNet50 relies on flower petals and green background features as shortcuts when classifying bee images [eval]
- IC-154CLIP ViT-L/14 text embeddings fail to capture fine-grained visual class similarities, ranking rottweiler and doberman at position 828 behind unrelated pairs [source]
- IC-1540CLS-token attention maps in pretrained ViT-t/16 exhibit high inter-layer correlation (cosine similarity up to 0.97) concentrated in layers 3–10, and MSA block outputs show high CKA in layers 2–8 [eval]
- IC-1543VGG19, ResNet50, ViT-Base, and DeiT-Base (ImageNet pretrained) achieve near-zero accuracy under query-based black-box attacks with 1000–10000 queries [eval]
- IC-193The OpenCLIP ResNet-50 model trained on CC12M contains an unintentional backdoor from birthday cake images in CC3M, achieving 98.92% attack success rate [eval]
- IC-194Temporal modeling in video models drives representational alignment to early visual cortex, while action classification task drives alignment to late brain areas [source]
- IC-258DNN accuracy on 3D perception tasks correlates with ImageNet object classification accuracy, suggesting 3D cues emerge as a byproduct of object recognition training [eval]
- IC-280CLIP, OpenCLIP, and SigLIP exhibit intra-modal misalignment: intra-modal similarity comparisons are suboptimal for image-to-image and text-to-text retrieval [eval]
- IC-281SLIP's intra-modal self-supervised loss reduces intra-modal misalignment, making inter-modal inversion unnecessary for image retrieval [eval]
- IC-290Zeroing out or doubling specific FFN neurons identified by the neuron path method causes significant accuracy changes in ViT and MAE models [eval]
- IC-291ViT-B/16 and MAE-B/16 exhibit nearly inverted distributions of knowledge neurons across layers despite identical architecture and training data [eval]
- IC-292Neuron paths in ViT-B/16 show class-specific neuron clustering and semantic similarity between image categories [eval]
- IC-293ViT-B/16 and ViT-B/32 are largely redundant: retaining only top-5 neurons per layer while zeroing all others preserves most classification accuracy [eval]
- IC-295The ViT model's ECE can be reduced to near-zero by trivial mean-replacement recalibration while maintaining test accuracy, but NLL increases from 65.35 to 144.66, demonstrating that ECE and accuracy alone are an insufficient reporting standard for calibration [eval]
- IC-318CLIP backbones from different architectures (ViTs and ResNets) trained with the same data and objective exhibit complementary strengths, with an oracle per-image backbone selection improving zero-shot accuracy by up to 43.5% over the best single backbone [eval]
- IC-319Different CLIP backbones exhibit distinct robustness profiles to specific image perturbations, with each architecture being most resilient to a different transformation [eval]
- IC-413Factors of variation in ImageNet-X are linearly decodable from the second-to-last-layer representations of ImageNet-pretrained ResNet50 and ViT-B/16 [source]
- IC-419CLIP ViT-B/16's layer-11 residual stream contains class-discriminative information in sparse SAE latent directions, and ablating class-specific top-k latents significantly degrades zero-shot classification accuracy [eval]
- IC-420CLIP ViT-B/16's SAE latent interpretability is depth-dependent: layer 11 encodes semantic object concepts while layers 2, 5, and 8 encode local shapes and attention patterns [train]
- IC-447CLIP ViT-B/16 produces noisy saliency maps and contains only 42 concept detectors, indicating poor visual interpretability [eval]
- IC-448CLIP ViT-L/14 achieves 0% accuracy under 2/255 and 4/255 L-infinity adversarial perturbations across all 15 evaluation datasets [eval]
- IC-449CLIP ViT-B/16 Grad-CAM explanations are highly sensitive to input noise, with SSIM dropping from 91.18% to 70.58% as noise standard deviation increases from 1/255 to 9/255 [eval]
- IC-542The PyTorch pretrained ResNet50 on ImageNet is vulnerable to (1, y)-ACE calibration attacks that increase ECE from 3.70% to 47.23% while preserving accuracy [eval]
- IC-627ResNet 18 and ViT B/16 retain significant ImageNet accuracy at 2–3 bit weight compression via JLCM [eval]
- IC-642CLIP can infer contextual attributes (orientation, illumination, etc.) from images with approximately 74% accuracy on a binary task [eval]
- IC-643Conditioning CLIP on correct contextual attributes in the text prompt improves zero-shot classification accuracy across 13 image transformations [eval]
- IC-646DINOv2, DeiT-III, and OpenCLIP repurpose approximately 2% of patch tokens in low-informative background areas as internal registers, discarding local patch information while aggregating global image information; DINO does not exhibit this behaviour [eval]
- IC-673Trained depthwise convolutional kernels in DS-CNN architectures converge to identifiable DoG-like patterns, with over 95% of ConvNeXtV2 and over 90% of ConvNeXt filters classifiable into a small set of clusters [source]
- IC-677CLIP ViT's image representation is primarily constructed by the last 4 MSA layers, with MLPs and early MSA layers contributing negligibly [eval]
- IC-678Specific attention heads in CLIP ViT-L's last 4 layers encode specific image properties (color, shape, location, counting, texture) that are linearly recoverable via text directions [eval]
- IC-684CLIP ViT-B/32 misclassifies 99% of forest satellite images as ocean when the word 'ocean' is overlaid as text [eval]
- IC-709Swin Transformer parameter redundancy is depth-dependent, with higher layers retaining fewer parameters than lower layers under module-aware pruning [eval]
- IC-732Off-the-shelf PyTorch ResNet classifiers are better calibrated than fine-tuned U-Net classifiers at high noise levels in the diffusion reverse process [eval]
- IC-737CLIP ViT-B/16 binarized dot products yield 0.50–0.58 accuracy on binary concept presence queries across five image classification datasets [eval]
- IC-738BLIP-2 ViT-G FlanT5XL achieves 0.70–0.87 zero-shot accuracy on binary concept presence queries, competitive on most datasets but weaker on fine-grained CUB-200 [eval]
- IC-739GPT-3.5-turbo-0613 combined with CLIP produces more faithful concept-salience pseudo-labels than LLaMA-2-13B-Chat, InstructBLIP, or LLaVA-1.5B on most of five datasets [eval]
- IC-935In pre-trained ViT, query vector Kruskal rank reaches the context size only after one self-attention layer, while general position fails at all depths [eval]
- IC-990ResNet-50 and DenseNet-101 exhibit a higher mean-to-variance ratio in penultimate pre-ReLU activations for in-distribution samples than for out-of-distribution samples, and the activation-based scaling factor is well-separated between ID and OOD [eval]