Light Dark anchor
Note anchor is the dataset's official page, which is the most authoritative reference that exists for it: the underlying reference is Krizhevsky's 2009 technical report "Learning multiple layers of features from tiny images", which carries no arXiv id or DOI Findings IC-028 SPADE, an abstaining classifier built on top of ResNet, ViT, and VGG models, detects out-of-distribution and adversarial samples with provable guarantees. [train] IC-038 Few embedding dimensions drive the modality gap in CLIP and SigLIP [eval] IC-077 VGG-16's loss landscape barrier height distinguishes adversarial from real inputs, enabling a detection method that outperforms baselines on DeepFool and C&W attacks [eval] IC-1017 Vision and language models pre-trained on noisy data exhibit degraded OOD transfer that is partially recoverable via SVD-based feature-space regularization [eval] IC-1105 PAC-Bayes generalization bounds for discrete class prompts on CLIP are within a few percentage points of the actual test error across CIFAR-10, CIFAR-100, ImageNet, FMOW, and OfficeHome [eval] IC-1106 CLIP prompts found by greedy search do not fit random labels: train and test error drop in tandem as the fraction of flipped labels increases, unlike a linear probe which achieves near-random accuracy [eval] IC-1140 Stable Diffusion's internal concept representations encode visual and structural similarities (shape, texture, color) that transcend textual semantics [eval] IC-152 CLIP's polysemantic neurons encode spurious correlations between unrelated concepts that can be exploited to generate adversarial misclassifications [eval] IC-1544 The latent spaces of pretrained foundational models across vision and text are not related by a single class of geometric transformations; the optimal alignment depends on the specific model pair, architecture, and dataset. [eval] IC-1579 ADM's noise prediction network exhibits exposure bias: during iterative sampling the l2-norm of its ε prediction is systematically larger than during training, and the sampling distribution variance exceeds the training variance with error accumulating toward the end of the chain [eval] IC-1602 ResNet18, ResNet34, and MobileNetV2 pre-trained on CIFAR10 have decision functions well-approximated by a kernel machine using the trace NTK, with Kendall-τ correlations of 0.776, 0.786, and 0.700 [eval] IC-1603 ResNet18's CIFAR10 classification decisions are driven by the bulk of training data rather than a sparse set of exemplars, as revealed by trntk data attribution [eval] IC-295 The ViT model's ECE can be reduced to near-zero by trivial mean-replacement recalibration while maintaining test accuracy, but NLL increases from 65.35 to 144.66, demonstrating that ECE and accuracy alone are an insufficient reporting standard for calibration [eval] IC-318 CLIP backbones from different architectures (ViTs and ResNets) trained with the same data and objective exhibit complementary strengths, with an oracle per-image backbone selection improving zero-shot accuracy by up to 43.5% over the best single backbone [eval] IC-330 Low local intrinsic dimension (LIDθ) of the learned manifold predicts memorization in Stable Diffusion v1.5, IDDPM, and StyleGAN2-ADA [eval] IC-448 CLIP ViT-L/14 achieves 0% accuracy under 2/255 and 4/255 L-infinity adversarial perturbations across all 15 evaluation datasets [eval] IC-548 CLIP-B/32 exhibits progressively increasing layer-wise representation similarity in both its vision encoder and text encoder, and the pattern also holds across modalities [eval] IC-604 CLIP ViT-B/32's CIFAR-10 image embeddings approximately satisfy a multi-cluster structure with near-orthogonal class-mean features [eval] IC-737 CLIP ViT-B/16 binarized dot products yield 0.50–0.58 accuracy on binary concept presence queries across five image classification datasets [eval] IC-738 BLIP-2 ViT-G FlanT5XL achieves 0.70–0.87 zero-shot accuracy on binary concept presence queries, competitive on most datasets but weaker on fine-grained CUB-200 [eval] IC-739 GPT-3.5-turbo-0613 combined with CLIP produces more faithful concept-salience pseudo-labels than LLaMA-2-13B-Chat, InstructBLIP, or LLaVA-1.5B on most of five datasets [eval] IC-917 ViT-S models show early-layer sensitivity to layer-wise averaging, with the averaging direction being far more disruptive than random perturbations of the same norm [eval] IC-990 ResNet-50 and DenseNet-101 exhibit a higher mean-to-variance ratio in penultimate pre-ReLU activations for in-distribution samples than for out-of-distribution samples, and the activation-based scaling factor is well-separated between ID and OOD [eval]