IC-128Released VLMs exhibit sycophancy, agreeing with incorrect user opinions while ignoring visual evidence, with LLaVA-1.5 showing the highest rate (94.6%) and InternLM-XComposer2-VL-1.8B the lowest (28.8%)
Shuo Li, Tao Ji, Xiaoran Fan, Linsheng Lu, Leyi Yang, Yuming Yang, Zhiheng Xi, Rui Zheng, Yuran Wang, xh.zhao, Tao Gui, Qi Zhang, Xuanjing Huang
The paper evaluates ten released VLMs on the MM-SY benchmark, where the model first answers a visual question correctly, then the user provides an incorrect opinion in a second round. The sycophancy rate measures how often the model abandons its correct answer. LLaVA-1.5 shows the highest sycophancy (99.4/94.6/89.7% across three tones), while InternLM-XComposer2-VL-1.8B shows the lowest (33.3/20.2/33.0%). Sycophancy varies by task (highest in object presence, lowest in object recognition) and by tone, but no single tone universally dominates. Multi-round persistence (up to 5 rounds) increases sycophancy by only about 5%.
Evaluation uses 150 questions per task from TDIUC; closed-source models (Gemini, GPT-4V) are evaluated via text matching rather than logits, which may affect comparability.