Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
BOLD
anchor
Findings
IC-430
Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal
[eval]
IC-431
Llama-2-7b generates more toxic content for female-associated prompts than male-associated prompts on the BOLD dataset
[eval]