MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl
anchor · artifact
- Note
- the artifact is the webdataset packaging of the captions split used in CLIP benchmarking, not the full COCO release
Findings
- FX-001Pairwise Banzhaf interactions explain CLIP similarity more faithfully than single-score methods [eval]
- IC-006The Retriever-Dictionary module improves object detection accuracy of YOLOv7, YOLOv9, Faster R-CNN, and Deformable DETR on COCO 2017 [eval]
- IC-018Object information is localized to specific visual tokens in LLaVA-1.5 [eval]
- IC-019Visual token representations in LLaVA-1.5 evolve to align with interpretable text tokens [eval]
- IC-020LLaVA-1.5 extracts object information directly from visual tokens to the last token in mid-late layers [eval]
- IC-023Local scaling, rank, and complexity of Stable Diffusion correlate with generation aesthetics, diversity, and memorization [eval]
- IC-038Few embedding dimensions drive the modality gap in CLIP and SigLIP [eval]
- IC-039Object bias in CLIP and SigLIP is not correlated with performance on attribute tasks [eval]
- IC-069EVA-CLIP's dense patch features are semantically contaminated by surrounding context, degrading their spatial quality [eval]
- IC-070Region-language alignment fine-tuning degrades EVA-CLIP's spatial awareness as measured by unsupervised segmentation [train]
- IC-106Logit lens on LLaVA and InstructBLIP image representations shows higher internal confidence for objects present in the image than for hallucinated objects [eval]
- IC-107Linear orthogonalization of LLaVA and InstructBLIP image features against text embeddings removes hallucinated objects at 83-86% individual rate versus 7-16% for correctly detected objects [eval]
- IC-1071Stable Diffusion v2.1's pixel-wise conditional mutual information localizes abstract words (adjectives, adverbs, verbs) more effectively than attention, but is less effective than attention for object segmentation [eval]
- IC-116In LLaVA-7B, multi-head attention modules drive hallucination more than MLP modules, and targeted intervention on specific hallucination heads reduces the hallucination rate by up to 1.7x [eval]
- IC-117Hallucination heads in LLaVA-7B and MiniGPT-4 are concentrated in the middle and deeper layers of the transformer [eval]
- IC-118Hallucination heads in LLaVA-7B and MiniGPT-4 allocate 4.75x more attention to text tokens than image tokens, and this pattern is inherited from the base language model [eval]
- IC-119The number of salient hallucination heads decreases as model size increases within the LLaVA family [eval]
- IC-1204SD-VAE 1.x and SD-VAE 2.x achieve lower reconstruction quality than the new SDXL-VAE on COCO 2017 [eval]
- IC-1405OpenFlamingo and Idefics models hallucinate objects not present in images, and increasing ICL shots beyond 4 amplifies hallucinations [eval]
- IC-148Language models represent semantically equivalent inputs from different data types (languages, code, images, audio) close together in intermediate layers, with the shared space scaffolded by the model's dominant language [eval]
- IC-280CLIP, OpenCLIP, and SigLIP exhibit intra-modal misalignment: intra-modal similarity comparisons are suboptimal for image-to-image and text-to-text retrieval [eval]
- IC-330Low local intrinsic dimension (LIDθ) of the learned manifold predicts memorization in Stable Diffusion v1.5, IDDPM, and StyleGAN2-ADA [eval]
- IC-334GPT-4V produces more poetic, emotion-focused image captions compared to Gemini-1.5-Flash's literal descriptions, with 99% model-matching accuracy on COCO [eval]
- IC-450LLaVA with the original CLIP encoder produces noisy, non-sparse attention maps that poorly localize to the objects described in generated text [eval]
- IC-457All four tested VLM decoders are heavily text-centric when generating answers, with text modality contributing 85-97% of the prediction signal [eval]
- IC-458Most VLM decoders show negative CC-SHAP on VALSE multiple-choice, indicating their explanations are less self-consistent than their answers, driven by a shift from text-dominant to image-dominant processing [eval]
- IC-479SAM's mask decoder exhibits attention drift to background or specific object parts under imprecise prompts, causing severe segmentation degradation [eval]
- IC-495All evaluated multimodal foundation models achieve average non-hallucination accuracy below 50% across six hallucination scenarios [source]
- IC-497Text-to-image models experience performance drops exceeding 10% under adversarial prompts, with spatial reasoning being the most vulnerable task across all models [source]
- IC-499ViT patch embeddings contain local semantic information beyond the [cls] token, as shown by performance degradation when restricting the output head to [cls] only or removing positional embeddings [eval]
- IC-582Instruction-tuned MLLMs (InstructBLIP, mPLUG-Owl, Idefics) achieve significantly better brain alignment than vision-only ViT-H and perform comparably to or better than CLIP-text across whole visual cortex and five visual ROIs [source]
- IC-633BLIP-2, LLaVA, and mPLUG-Owl show a trade-off between caption length and hallucination rate on COCO [eval]
- IC-647DINOv2's feature-map artifacts cause it to be incompatible with the LOSt unsupervised object discovery method, scoring far below DINO [eval]
- IC-667CLIP's intermediate layer features encode object boundaries recoverable by k-means clustering, a property absent in shallow and deep layers [eval]
- IC-668SAM's edge-oriented segmentation yields high recall but very low precision because it cannot distinguish object boundaries from interior edges [eval]
- IC-669DINOv2's features, when clustered, produce smooth semantic regions but lack instance-level boundary delineation [eval]
- IC-763CLIP and OpenCLIP fall short of human discriminative accuracy on vision tasks, with performance dropping substantially under hard negatives [eval]
- IC-765BLIP-2, BLIP, InstructBLIP, Bard, and BingChat fall short of human accuracy in answering questions about Midjourney-generated images [eval]
- IC-932Pre-trained DNN object detectors show a sharp falloff in peripheral detection performance with increasing eccentricity, degrading to near-chance by 20°, while human performance degrades gradually [source]
- IC-933DNN object detectors do not exhibit the same sensitivity to image clutter as humans in peripheral object detection [source]
- IC-934DNN object detectors do not exhibit the same object size effect as humans in peripheral object detection [source]