IC-244No LLM demonstrates consistently strong fairness across both comprehension-focused and bias-resistance multi-turn tasks; models show complementary failure patterns

Zhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu Liu

SourceFairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs

Across six LLMs evaluated on six multi-turn fairness tasks, the paper finds that models excel at different task types. Llama-2-Chat (7B, 13B) performs poorly on comprehension tasks like anaphora ellipsis (14.93% and 18.35% bias) but is less affected by interaction interference like negative feedback (2.75% and 2.89%). Conversely, Mistral-7B-Instruct handles scattered questions better (11.55%) but is highly susceptible to interference from misinformation (58.10%). The paper concludes that no model has yet demonstrated consistently strong fairness across both comprehension-focused and bias-resistance tasks.

Evidence
correlational
Key metric
Llama-2-7B-Chat: anaphora ellipsis 14.93%, negative feedback 2.75%; Mistral-7B-IT: scattered questions 11.55%, interference misinformation 58.10%; Llama-3.1-8B-IT: scattered questions 13.56%, interference misinformation 51.31%
Caveat
The paper attributes the differences to 'variations in alignment paradigms and instruction-following capabilities' but does not isolate the causal mechanism; the pattern is correlational across models.
Model
Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b, Llama-2-13B-Chat, Llama 3.1 8B Instruct, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct-v0.3, Gemma Gemma-7B-IT
Concepts
Failure mode
Datasets
FairMT-10K [eval], MT-Bench-101 [eval]
Methods
Llama-Guard-3-8B [eval]
Related findings
IC-242, IC-243
Extraction
automatic-extraction