The paper evaluates the effect of LASER on the performance of Phi-3 and Llama-3.1-8B models on the GSM8K math word problem benchmark using few-shot Chain-of-Thought prompting. For 1-shot and 2-shot settings, applying LASER improves accuracy compared to the full model. For example, Phi-3 accuracy improves from 56.0% to 66.1% in the 1-shot setting.
Evidence
correlational
Key metric
Phi-3 1-shot accuracy improves from 56.0% to 66.1% with LASER. Llama-3.1-8B 1-shot accuracy improves from 44.7% to 46.1% with LASER.
Caveat
In the standard 8-shot setting, performance is worse with LASER.