IC-017Truncating MLP weights improves few-shot Chain-of-Thought reasoning accuracy on GSM8K for Phi-3 and Llama-3.1-8B

Lei Chen, Joan Bruna, Alberto Bietti

SourceDistributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers

The paper evaluates the effect of LASER on the performance of Phi-3 and Llama-3.1-8B models on the GSM8K math word problem benchmark using few-shot Chain-of-Thought prompting. For 1-shot and 2-shot settings, applying LASER improves accuracy compared to the full model. For example, Phi-3 accuracy improves from 56.0% to 66.1% in the 1-shot setting.

Evidence
correlational
Key metric
Phi-3 1-shot accuracy improves from 56.0% to 66.1% with LASER. Llama-3.1-8B 1-shot accuracy improves from 44.7% to 46.1% with LASER.
Caveat
In the standard 8-shot setting, performance is worse with LASER.
Model
Phi-3, Llama 3.1 8B
Datasets
GSM8K [eval]
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [primary], Layer-Selective Rank Reduction / LASER [primary]
Extraction
automatic-extraction