SourceMeasuring Vision-Language STEM Skills of Neural Models
The paper plots the relationship between CLIP's softmax confidence and its actual accuracy on the STEM test set. The zero-shot model is overconfident: its predicted confidence is systematically higher than its realized accuracy, and the two are only loosely correlated. After fine-tuning on the STEM training split, the model becomes more calibrated, suggesting the miscalibration is partly due to the zero-shot transfer setting rather than an inherent architectural property.