IC-031Empowered persona prompts and reflection mechanisms reduce conformity in large language models

Zhiyuan Weng, Guikun Chen, Wenguan Wang

SourceDo as We Do, Not as You Think: the Conformity of Large Language Models

The paper tests two prompt-based mitigation strategies: empowered persona prompts (encouraging independent thinking) and reflection prompts (asking models to re-evaluate answers). For Llama3-70B, empowered persona increases IR from 28.6% to 40.0% and reduces CRD from 69.9% to around 35-55% across variants; reflection increases IR from 28.6% to 68.5% and reduces CRD from 69.9% to 35.2%. For Qwen2-72B, empowered persona increases IR from 57.6% to 68.6%, while reflection produces mixed results (CRT increases from 30.5% to 45.0%) suggesting the strategy interacts with model characteristics.

Evidence
interventional
Key metric
Llama3-70B: IR from 28.6% to 40.0% (empowered persona) and to 68.5% (reflection). Qwen2-72B: IR from 57.6% to 68.6% (empowered persona); CRT increases from 30.5% to 45.0% under reflection while CRC, CRW, CRD decrease. Prompt P3 and P5 selected as best performing.
Caveat
Effectiveness depends on carefully calibrated prompts and may vary across models; universal prompts are challenging to design. Reflection increased conformity for Qwen2-72B on trust protocol, indicating potential backfire. Strategies' generalizability across diverse LLM characteristics remains a challenge.
Model
Llama 3, Qwen 2
Datasets
BENCHFORM [eval]
Extraction
automatic-extraction