IC-243Bias ratio on certain multi-turn fairness tasks decreases with model size in the Gemma-2 and Qwen2.5 families

Zhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu Liu

SourceFairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs

On the more challenging FairMT-1K subset, the paper tests 15 LLMs including three sizes each of Gemma-2 and Qwen2.5. For tasks such as scattered questions, fixed format, and negative feedback, the proportion of biased responses tends to decrease as model size increases. In the Gemma-2 family, scattered questions bias drops from 83.03% (2B) to 20.00% (27B), and fixed format drops from 51.52% (2B) to 6.67% (27B). The authors attribute this to increased parameters enhancing comprehension of user instructions in multi-turn dialogues.

Evidence
correlational
Key metric
Gemma-2 scattered questions: 83.03% (2B), 26.06% (9B), 20.00% (27B); fixed format: 51.52% (2B), 39.39% (9B), 6.67% (27B); negative feedback: 23.03% (2B), 40.61% (9B), 9.70% (27B)
Caveat
The trend is not strictly monotonic for all tasks and families; Qwen2.5 shows less consistent decrease (e.g., negative feedback: 93.33% at 0.5B, 95.76% at 3B, 87.27% at 7B). The paper states the proportion 'tends to decrease' rather than guaranteeing it.
Model
Gemma 2 2B It, Gemma-2-9B-IT, Gemma-2-27B-IT, Qwen2.5 Qwen2.5-0.5B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct
Concepts
Scale-dependent behaviour
Related findings
IC-242, IC-244
Extraction
automatic-extraction