The paper demonstrates that vision-language models also suffer from position bias. For Fuyu-8B, the loss on the ground-truth token is consistently lower when a real-world image is inserted at the bottom of a large black background rather than at other positions. For GPT-4V, the model correctly identifies the M110 satellite galaxy when it appears at the top of the Andromeda galaxy image (flipped versions c and d) but incorrectly identifies it as M32 when it appears at the bottom (original and one flip, versions a and b).
Evidence
correlational
Caveat
The GPT-4V result is based on a single image with four flips (qualitative, n=4). The Fuyu-8B result is shown as a loss curve in Figure 1 without specific numerical values printed in the text.