IC-1499Self-rationalization quality and task accuracy scale with model size across GPT-3, FLAN-T5, and LLaMA on five QA datasets

Sahana Ramnath, Brihi Joshi, Skyler Hallinan, Ximing Lu, Liunian Harold Li, Aaron Chan, Jack Hessel, Yejin Choi, Xiang Ren

SourceTailoring Self-Rationalizers with Multi-Reward Distillation

The paper evaluates GPT-3 (175B), FLAN-T5 (L, XL, XXL), and LLaMA (7B, 65B) on self-rationalization across StrategyQA, QuAREL, OpenBookQA, Numersense, and QASC using few-shot chain-of-thought prompting. Task accuracy and rationale quality metrics (plausibility, consistency, diversity) generally improve with model scale: GPT-3 achieves the highest average NRG on most datasets, FLAN-T5-XXL and LLaMA-65B are competitive, while LLaMA-7B and FLAN-T5-L show substantially lower performance, particularly on Numersense (17.59% and 26.13% accuracy respectively) and QASC (24.19% and 61.02%).

Evidence
correlational
Key metric
Task accuracy: GPT-3 69.0/83.33/85.94/74.37/80.24; FLAN-T5-XXL 70.52/77.54/80.32/61.81/74.84; LLaMA-65B 72.27/76.27/73.30/36.18/75.59; LLaMA-7B 59.17/56.70/40.76/17.59/24.19; FLAN-T5-L 54.59/77.36/60.64/26.13/61.02 (StrategyQA/QuAREL/OpenBookQA/Numersense/QASC). Avg NRG: GPT-3 72.13/79.46/80.36/80.84/80.31; LLaMA-7B 67.29/66.18/63.07/59.40/53.05
Caveat
All reference LLMs are evaluated via few-shot prompting with the same demonstrations; no fine-tuning is applied to them. The paper notes that MARIO (0.7B) still lags behind GPT-3 on plausibility and task accuracy for most datasets.
Model
GPT-3 / GPT base, FLAN-T5, LLaMA
Concepts
Scale-dependent behaviour
Datasets
StrategyQA [eval], QuAREL [eval], OpenBookQA [eval], Numersense [eval], QASC [eval]
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [eval]
Extraction
automatic-extraction