The paper evaluates calibration (ECE, Brier score) of LLaMA-7B, LLaMA-13B, LLaMA-30B, LLaMA2-7B, and LLaMA2-13B across phrase-, sentence-, and paragraph-level generation tasks. On phrase-level tasks, the larger models in each family achieve lower ECE and Brier scores than their smaller counterparts. However, on sentence- and paragraph-level tasks, the scaling trend breaks down: for example, LLaMA-30B has higher ECE than LLaMA-7B on WikiQA (0.142 vs. 0.108) and WikiGen (0.165 vs. 0.102), indicating that the within-family scaling benefit does not extend to longer generations.
The scaling trend is inconsistent across individual tasks within the same length category; the paper states the effect 'may not hold true' rather than being uniformly absent.