IC-736Vicuna-13B outperforms Vicuna-7B on factual knowledge tasks by 5.4% on average

Xuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo, Lijie Wen, Philip S. Yu, Zhijiang Guo

SourceTowards Understanding Factual Knowledge of Large Language Models

The paper compares Vicuna-7B and Vicuna-13B across all Pinocchio tasks under few-shot CoT prompting. As model parameters increase from 7B to 13B, performance on factual questions improves correspondingly, with an average increase of 5.4%. The authors interpret this as evidence that LLMs with more parameters can store more world knowledge and have stronger factual knowledge recognition capabilities.

Evidence
correlational
Key metric
average increase of 5.4% from Vicuna-7B to Vicuna-13B; few-shot CoT overall: Vicuna-7B acc 48.5 F1 40.6, Vicuna-13B acc 47.0 F1 42.5
Caveat
The paper notes that due to limited computing resources, hyperparameter exploration was only done for these two Vicuna variants. The 5.4% figure is an average across tasks; individual task differences vary.
Model
Vicuna Vicuna-7B, Vicuna-13B
Concepts
Scale-dependent behaviour
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [eval]
Related findings
IC-733, IC-734, IC-735
Extraction
automatic-extraction