IC-368Larger LMs (GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro) exhibit better calibration than their smaller counterparts (GPT-4o-mini, Claude-3-Haiku, Gemini-1.5-Flash) when verbalizing confidence with certainty phrases

Peiqi Wang, Barbara D. Lam, Yingcheng Liu, Ameneh Asgari-Targhi, Rameswar Panda, William M Wells, Tina Kapur, Polina Golland

SourceCalibrating Expressions of Certainty

The paper evaluates six released LMs on SCIQ and TruthfulQA, prompting them to answer questions and express confidence by selecting from a fixed set of 12 certainty phrases. Calibration is measured via ECE using a distribution-based estimator. The paper reports that larger models consistently outperform their smaller variants in calibration: GPT-4o achieves ECE 0.07 on SCIQ and 0.22 on TruthfulQA, while the paper states that GPT-4o-mini and other smaller variants show worse calibration. This pattern holds across all three model families tested.

Evidence
correlational
Key metric
GPT-4o phrase uncalibrated: SCIQ ece 0.07, TruthfulQA ece 0.22; Claude-3.5-Sonnet: SCIQ ece 0.06, TruthfulQA ece 0.25; Gemini-1.5-Pro: SCIQ ece 0.17, TruthfulQA ece 0.28
Caveat
Specific ECE numbers for the smaller variants (GPT-4o-mini, Claude-3-Haiku, Gemini-1.5-Flash) are shown in Figure 4 but not printed as text values in the paper body.
Model
GPT-4o mini, Claude 3.5 Sonnet, Claude 3 Haiku, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Gemini 1.5 Flash
Concepts
Scale-dependent behaviour
Datasets
SciQ / SciQA [eval], TruthfulQA / TruthfulQA MC1 [eval]
Methods
Platt scaling [compared-to]
Related work
Tian et al. 2023 (Just Ask for Calibration) [context]
Related findings
IC-369
Extraction
automatic-extraction