IC-353Llama-3.1-instruct 8b and 70b fail the harder NIAH test (sandwich needle) but pass the easier passkey retrieval test

Peng Xu, Wei Ping, Xianchao Wu, Chejian Xu, Zihan Liu, Mohammad Shoeybi, Bryan Catanzaro

SourceChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities

The paper evaluates Llama-3.1-instruct 8b and 70b on two variants of the Needle-in-a-Haystack test up to 128k context. With the harder needle ('the best thing to do in san francisco is eat a sandwich and sit in dolores park on a sunny day'), neither model achieves reliable retrieval across depths and positions. With the easier passkey needle ('the pass key is 385243. remember it. 385243 is the pass key'), both models pass. This indicates a threshold effect where the models can retrieve short, distinctive tokens but fail on longer, less distinctive sentences.

Evidence
observational
Caveat
The NIAH test is a synthetic evaluation; the authors note it 'does not accurately represent real-world downstream task performance.' Results are shown in figures rather than tabulated numbers.
Model
Llama 3.1 8B Instruct, 70B Instruct
Concepts
Failure mode
Datasets
Needle-in-a-haystack [eval]
Related findings
IC-351, IC-352
Extraction
automatic-extraction