Light Dark The model's response to the same evidence depends on where in the input that evidence sits. Distinct from an attribution artefact: the bias is in the model's own output, not in the map drawn to explain it, so it needs the evidence to be moved and the output remeasured.
Findings IC-1014 A single adversarial token embedding appended to any input prompt overwrites the prompt in Stable Diffusion v2.1 to generate a target object, with CLIP similarity to the original prompt (0.742) remaining higher than to the target (0.546) IC-1090 ICL in LLaMA, LLaMA-2, and Falcon models preferentially uses in-context label information closer to the query rather than treating all examples equally IC-1135 LLMs show order sensitivity to evidence position in context, with PaLM2 and LLaMA2-7B showing memorization ratio variations exceeding 30% IC-1148 ChatGPT and InstructGPT exhibit positional bias in their reliance on in-context examples, with ChatGPT showing decreasing attention by position and InstructGPT showing a U-shaped pattern IC-1373 All five evaluated LLMs show a strong positional bias in constrained text generation, with first-position constraints nearly always satisfied but last- and arbitrary-position constraints causing major performance drops IC-1585 LLaMA-7B and GPT-J-6B exhibit positional bias in instruction following: zero-shot performance varies significantly when the instruction is moved from after to before the context IC-180 GPT-4o-mini's resistance to incorrect context depends on the position of the context relative to the question in the prompt IC-238 LLMs fail to follow user preferences in zero-shot settings, with accuracy below 10% at 10 turns and near zero at 300 turns IC-241 Preference following degrades when the preference is placed in the middle of a long conversation, extending the lost-in-the-middle effect to preference tracking IC-245 Pretrained LLMs produce duration-dependent outputs that are incompatible with a discrete token interpretation IC-256 GPT-2 small's attention product functions p_i^T k^T q p_j are approximately translation-invariant across all 144 heads IC-305 Mistral-instruct-7b's susceptibility to prompt-injected data extraction follows a U-shaped curve depending on the position of the adversarial prompt within the context window IC-341 The effectiveness of GCG and PAIR attacks on Llama-2-7b-chat is sensitive to the order of adversarial tokens, with swapping the two halves of the GCG suffix reducing the created high-importance region by 23% IC-367 Sparsifying initial tokens of the prefill phase causes disproportionate degradation in Llama-3-8B due to attention sink behavior IC-391 Llama-3 and Qwen-1.5 models exhibit position bias in LM-as-a-judge, retrieval-augmented QA, and math reasoning, with larger models showing less bias IC-392 Fuyu-8B and GPT-4V exhibit position bias in visual recognition, with model performance depending on where the target object appears in the image IC-476 GPT-4 and GPT-3.5 show strong positional bias in Chinese idiom character completion IC-579 ICL hidden states exhibit positional bias: representations of the same input are more similar when the input appears at similar positions in the sequence IC-588 Topologically ordered context improves relational reasoning over random ordering across nearly all LLMs IC-602 Lightweight LLMs exhibit positional bias in sequential checklist judgments, with judgment inconsistency increasing as the position of the item in the multi-turn dialogue grows IC-681 GPT-3.5 exhibits positional bias when judging which of two LLM responses is superior IC-945 LLaMA-2, MPT, Falcon, Pythia, and BERT-base-uncased allocate disproportionate attention to initial tokens regardless of their semantic content IC-998 GPT-3.5's zero-shot accuracy on OGBN-ARXIV depends on the position of the title relative to the abstract in the prompt: 0.720 when abstract precedes title, 0.695 when title precedes abstract SY-002 Sybil responds more weakly to nodules near the pleura, where adenocarcinoma tends to appear