IC-351GPT-4-turbo-2024-04-09, Qwen2-72b-instruct, and Llama-3.1-70b-instruct show distinct performance profiles across ultra-long, 32k, and 4k context benchmarks

Peng Xu, Wei Ping, Xianchao Wu, Chejian Xu, Zihan Liu, Mohammad Shoeybi, Bryan Catanzaro

SourceChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities

The paper evaluates several released long-context LLMs on three benchmark suites covering different context lengths. On ultra-long (>100k) InfiniteBench tasks, GPT-4-turbo-2024-04-09 scores 33.16, Qwen2-72b-instruct 39.77, and Llama-3.1-70b-instruct 39.81. On 32k tasks, GPT-4-turbo-2024-04-09 leads at 51.93, ahead of Qwen2-72b-instruct (49.94) and Llama-3.1-70b-instruct (49.92). On 4k ChatRAG tasks, GPT-4-turbo-2024-04-09 scores 54.72, Qwen2-72b-instruct 54.06, and Llama-3.1-70b-instruct 52.12. The ranking shifts across context lengths: GPT-4-turbo is strongest at 32k but weakest at ultra-long among the top three.

Evidence
correlational
Key metric
GPT-4-turbo-2024-04-09: 33.16 (ultra-long), 51.93 (32k), 54.72 (4k); Qwen2-72b-instruct: 39.77, 49.94, 54.06; Llama-3.1-70b-instruct: 39.81, 49.92, 52.12
Caveat
The authors note that Qwen2-72b-instruct and Llama-3.1-70b-instruct used extensive 32k pretraining and large SFT data, while the authors used a smaller continued pretraining corpus, which may explain the 32k gap.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4-turbo-20240409, Qwen 2 Qwen2-72B-Instruct, Llama 3.1 70B Instruct, 8B Instruct, Yi 34B 200K, Llama 3 Llama-3-70B-Instruct-Gradient-262K, Llama3-ChatQA-1.5-70B
Datasets
LongBench [eval], Scrolls [eval]
Methods
E5-Mistral [supporting]
Related findings
IC-352, IC-353
Extraction
automatic-extraction