IC-097Chain-of-thought prompting outperforms few-shot direct QA for Llama 3 models above 8B parameters, while few-shot is best below 3B
Ricardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe, Stefan Bechtold, Christoph Engel, Jens Frankenreiter, Krishna P. Gummadi, Moritz Hardt, Michael Livermore
The paper compares three prompting strategies (zero-shot direct QA, 15-shot direct QA, and zero-shot chain-of-thought) across the Llama 3 instruct family. For models below 3B parameters, few-shot direct QA performs best. For models above 8B parameters, chain-of-thought is superior. The largest model evaluated with few-shot, Llama 3.3 70B Instruct, does not benefit from including in-context examples at all. CoT yields modest average improvements of 2-3 accuracy points for both the 8B and 70B models.
CoT evaluation was limited to Llama 3 8B and 70B Instruct due to compute cost (over 500 H100 GPU hours). Few-shot evaluation used 15 examples, constrained by the 128k context window and 8k token opinions.