The paper compares Vicuna-7B and Vicuna-13B across all Pinocchio tasks under few-shot CoT prompting. As model parameters increase from 7B to 13B, performance on factual questions improves correspondingly, with an average increase of 5.4%. The authors interpret this as evidence that LLMs with more parameters can store more world knowledge and have stronger factual knowledge recognition capabilities.
Evidence
correlational
Key metric
average increase of 5.4% from Vicuna-7B to Vicuna-13B; few-shot CoT overall: Vicuna-7B acc 48.5 F1 40.6, Vicuna-13B acc 47.0 F1 42.5
Caveat
The paper notes that due to limited computing resources, hyperparameter exploration was only done for these two Vicuna variants. The 5.4% figure is an average across tasks; individual task differences vary.