IC-188LLMs show constraint-type-specific performance on system message following, with weaker models exhibiting large variance across constraint categories

Yanzhao Qin, Tao Zhang, Yanjun Shen, Wenjing Luo, sunhaoze, Yan Zhang, Yujing Qiao, weipeng chen, Zenan Zhou, Wentao Zhang, Bin CUI

SourceSysBench: Can LLMs Follow System Message?

The paper evaluates 16 LLMs on 500 system messages containing six constraint types (action, content, background, role, format, style). GPT-4o leads across all types with an overall CSR of 87.1%. Stronger models show relatively uniform performance across constraint types, while weaker models show significant variance. For example, Qwen2-7b achieves 81.0% on role constraints but drops to 43.3% on style constraints. ERNIE-4, despite performing well on other instruction-following benchmarks, scores only 50.7% overall on system message following, comparable to Qwen2-7b.

Evidence
correlational
Key metric
GPT-4o CSR 87.1% (action 86.8%, content 86.9%, background 87.2%, role 93.5%, format 87.4%, style 86.5%); Qwen2-7b role 81.0% vs. style 43.3%; ERNIE-4 total 50.7%
Caveat
Evaluation uses GPT-4o as a model-based verifier; the paper reports 97.3% human-model consistency at constraint granularity but this is a proxy for ground truth.
Model
GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4-turbo-20240409, Claude 3 Opus, Llama 3.1 70B Instruct, 8B Instruct, Mixtral 8x22B Instruct, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo-20231106, Qwen 2.5 72B Instruct, Qwen 2 Qwen2-72B-Instruct, 7B Instruct, GLM-4 GLM-4-0520, GLM-4-9B-Chat, DeepSeek-V2-0628, Moonshot-v1-8k
Concepts
Failure mode
Related work
FollowBench [compared-to], IFEval / IFEval-Simple [compared-to]
Related findings
IC-189, IC-190, IC-191
Extraction
automatic-extraction