The paper evaluates six released LMs on SCIQ and TruthfulQA, prompting them to answer questions and express confidence by selecting from a fixed set of 12 certainty phrases. Calibration is measured via ECE using a distribution-based estimator. The paper reports that larger models consistently outperform their smaller variants in calibration: GPT-4o achieves ECE 0.07 on SCIQ and 0.22 on TruthfulQA, while the paper states that GPT-4o-mini and other smaller variants show worse calibration. This pattern holds across all three model families tested.
Specific ECE numbers for the smaller variants (GPT-4o-mini, Claude-3-Haiku, Gemini-1.5-Flash) are shown in Figure 4 but not printed as text values in the paper body.