Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Bias in Bios
anchor
Findings
IC-1137
Stable Diffusion represents certain concepts primarily through specific named exemplars rather than abstract category features
[eval]
IC-1139
Stable Diffusion encodes social biases in its internal concept representations that are not always visually apparent
[eval]
IC-1584
LLaMA-7B and GPT-J-6B fail to interpret textual emphasis markers, with marked prompting degrading performance substantially
[eval]
IC-1585
LLaMA-7B and GPT-J-6B exhibit positional bias in instruction following: zero-shot performance varies significantly when the instruction is moved from after to before the context
[eval]
IC-1586
In LLaMA-7B, steering all attention heads degrades JSON format accuracy below zero-shot, while steering a subset of 50-100 heads selected via multi-task profiling raises it to 96.64; performance varies dramatically across the 32 layers and individual heads
[eval]
IC-161
Linear probes on Pythia-70m and Gemma-2-2b trained on the ambiguous Bias in Bios set rely on gender as a spurious feature, with gender accuracy far exceeding profession accuracy
[eval]
IC-1630
Llama and Pythia models represent entity-attribute bindings via additive binding id vectors that form a continuous subspace with metric structure
[eval]
IC-1631
Binding id mechanism fidelity increases with model size in both Llama and Pythia families
[eval]