Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-351
GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct show distinct performance profiles across ultra-long, 32k, and 4k context benchmarks
IC-352
RAG with sufficient retrieved tokens outperforms direct long-context for Qwen2-72b-instruct on >100k tasks, while at 32k the default RAG setting underperforms direct long-context for GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct
IC-353
Llama-3.1-instruct 8b and 70b fail the harder NIAH test (sandwich needle) but pass the easier passkey retrieval test