IC-602Lightweight LLMs exhibit positional bias in sequential checklist judgments, with judgment inconsistency increasing as the position of the item in the multi-turn dialogue grows

Tianjun Wei, Wei Wen, Ruizhi Qiao, Xing Sun, Jianghong Ma

SourceRocketEval: Efficient automated LLM evaluation via grading checklist

The paper transforms WildBench checklist items into a multi-turn dialogue format and measures how the disagreement ratio on the current item's judgment changes as a function of the item's position in the sequence. For all five tested models (Llama-3-70B, Llama-3-8B, Qwen2-72B, Qwen2-7B, Qwen2-1.5B), the proportion of inconsistencies shows a growing trend as the number of previous questions increases. The Llama-3 series shows overall higher consistency compared to the Qwen-2 series. The paper identifies this positional bias as a second key mechanism (alongside high uncertainty) that degrades lightweight LLMs' reliability as sequential evaluators.

Evidence
correlational
Key metric
Disagreement ratio increases with position for all models; Llama-3 series shows higher consistency than Qwen-2 series (Figure 5, positions 1-7)
Caveat
The specific disagreement ratios per position are shown only in Figure 5 and not reported as exact numbers in the text. The effect is measured on WildBench checklist items in a multi-turn dialogue format.
Model
Qwen 2 Qwen2-1.5B, 7B, 72B, Llama 3 8B, 70B
Concepts
Positional bias
Datasets
WildBench [eval]
Related work
Split and Merge: Aligning Position Biases in LLM-based Evaluators [context]
Related findings
IC-601
Extraction
automatic-extraction