Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Large Language Models can Become Strong Self-Detoxifiers
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-430
Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal
IC-431
Llama-2-7b generates more toxic content for female-associated prompts than male-associated prompts on the BOLD dataset