IC-392Fuyu-8B and GPT-4V exhibit position bias in visual recognition, with model performance depending on where the target object appears in the image

Ziqi Wang, Hanlin Zhang, Xiner Li, Kuan-Hao Huang, Chi Han, Shuiwang Ji, Sham M. Kakade, Hao Peng, Heng Ji

SourceEliminating Position Bias of Language Models: A Mechanistic Approach

The paper demonstrates that vision-language models also suffer from position bias. For Fuyu-8B, the loss on the ground-truth token is consistently lower when a real-world image is inserted at the bottom of a large black background rather than at other positions. For GPT-4V, the model correctly identifies the M110 satellite galaxy when it appears at the top of the Andromeda galaxy image (flipped versions c and d) but incorrectly identifies it as M32 when it appears at the bottom (original and one flip, versions a and b).

Evidence
correlational
Caveat
The GPT-4V result is based on a single image with four flips (qualitative, n=4). The Fuyu-8B result is shown as a loss curve in Figure 1 without specific numerical values printed in the text.
Model
Fuyu Fuyu-8B, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision
Concepts
Positional bias
Related findings
IC-391
Extraction
automatic-extraction