Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
RealToxicityPrompts
anchor
Findings
IC-1007
LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedback
[eval]
IC-430
Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal
[eval]
IC-462
GPT-2 encodes toxicity in a low-dimensional linear subspace of its MLP layers, concentrated in higher layers
[eval]
IC-463
DPO's first-step gradients in GPT-2 are correlated with the toxic subspace, with stronger alignment in later layers and with more samples
[eval]
IC-815
RLHF on general-purpose preference data increases stereotypical bias and decreases truthfulness in Pythia and Llama-7B models
[eval]
IC-818
RLHF on general-purpose preference data has negligible net effect on toxicity in Pythia and Llama-7B models
[eval]