Across all phrase- and sentence-level tasks in the CAT benchmark, GPT-2 XL (1.5B) achieves consistently lower ECE and Brier scores than all larger models tested, including GPT-J (6B), LLaMA-7B/13B/30B, LLaMA2-7B/13B, and Vicuna-13B. For example, on NQ, GPT-2 XL has ECE 0.045 versus LLaMA-30B at 0.169. The authors attribute this to GPT-2 XL and GPT-J rarely exhibiting overconfidence on incorrect generations. This shows that model size alone does not determine calibration quality across families.