Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Jigsaw Unintended Bias in Toxicity Classification
anchor
Findings
IC-430
Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal
[source]