Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Number Cookbook: Number Understanding of Language Models and How to Improve It
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-112
Released LLMs show a reproducible failure mode where numerical task accuracy degrades sharply as input digit length increases
IC-113
Released LLMs show a reproducible failure mode where accuracy on fraction and scientific notation tasks falls below 20% even for the shortest inputs
IC-114
Released LLMs cannot reliably identify a specific digit in a number as the number's length increases, with GPT-4o achieving only 20% on get-digit in the xl range
IC-115
NUPA performance is largely independent of model size within a family: GPT-4o and GPT-4o-mini show nearly identical performance, as do Qwen2-72B and Qwen2-7B