IC-602Lightweight LLMs exhibit positional bias in sequential checklist judgments, with judgment inconsistency increasing as the position of the item in the multi-turn dialogue grows
Tianjun Wei, Wei Wen, Ruizhi Qiao, Xing Sun, Jianghong Ma
The paper transforms WildBench checklist items into a multi-turn dialogue format and measures how the disagreement ratio on the current item's judgment changes as a function of the item's position in the sequence. For all five tested models (Llama-3-70B, Llama-3-8B, Qwen2-72B, Qwen2-7B, Qwen2-1.5B), the proportion of inconsistencies shows a growing trend as the number of previous questions increases. The Llama-3 series shows overall higher consistency compared to the Qwen-2 series. The paper identifies this positional bias as a second key mechanism (alongside high uncertainty) that degrades lightweight LLMs' reliability as sequential evaluators.
Evidence
correlational
Key metric
Disagreement ratio increases with position for all models; Llama-3 series shows higher consistency than Qwen-2 series (Figure 5, positions 1-7)
Caveat
The specific disagreement ratios per position are shown only in Figure 5 and not reported as exact numbers in the text. The effect is measured on WildBench checklist items in a multi-turn dialogue format.