IC-470LOVT and MGCA, which use implicit cross-attention local alignment, show only marginal improvement over CLIP in 3D CT diagnosis (AUC 69.4 and 70.1 vs 68.4)
Zhongyi Shui, Jianpeng Zhang, Weiwei Cao, Sinuo Wang, Ruizhe Guo, Le Lu, Lin Yang, Xianghua Ye, Tingbo Liang, Qi Zhang, Ling Zhang
The paper evaluates LOVT and MGCA, two methods that integrate global contrastive learning with implicit local alignment via cross-attention, on the same 54 zero-shot CT diagnosis tasks. LOVT achieves AUC 69.4 and MGCA achieves AUC 70.1, only marginally above CLIP's 68.4. The authors interpret this as evidence that implicit local alignment, which works for 2D chest X-rays, is insufficient for 3D CT volumes that encompass hundreds of anatomical structures and far more complex reports.
The comparison is on the authors' own dataset; the 'marginal improvement' is only 1-2 AUC points, and the authors do not ablate whether the cross-attention mechanism itself is the bottleneck versus other factors.