IC-470LOVT and MGCA, which use implicit cross-attention local alignment, show only marginal improvement over CLIP in 3D CT diagnosis (AUC 69.4 and 70.1 vs 68.4)

Zhongyi Shui, Jianpeng Zhang, Weiwei Cao, Sinuo Wang, Ruizhe Guo, Le Lu, Lin Yang, Xianghua Ye, Tingbo Liang, Qi Zhang, Ling Zhang

SourceLarge-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding

The paper evaluates LOVT and MGCA, two methods that integrate global contrastive learning with implicit local alignment via cross-attention, on the same 54 zero-shot CT diagnosis tasks. LOVT achieves AUC 69.4 and MGCA achieves AUC 70.1, only marginally above CLIP's 68.4. The authors interpret this as evidence that implicit local alignment, which works for 2D chest X-rays, is insufficient for 3D CT volumes that encompass hundreds of anatomical structures and far more complex reports.

Evidence
correlational
Key metric
LOVT AUC 69.4, MGCA AUC 70.1, CLIP AUC 68.4 on 54 zero-shot diagnosis tasks (Table 1)
Caveat
The comparison is on the authors' own dataset; the 'marginal improvement' is only 1-2 AUC points, and the authors do not ablate whether the cross-attention mechanism itself is the bottleneck versus other factors.
Model
LOVT, MGCA, CLIP / CLIP-ViT (LC)
Concepts
Failure mode
Datasets
MedVL-CT69K [eval]
Related work
LOVT [compared-to], MGCA [compared-to]
Related findings
IC-469
Extraction
automatic-extraction