Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-462
GPT-2 encodes toxicity in a low-dimensional linear subspace of its MLP layers, concentrated in higher layers
IC-463
DPO's first-step gradients in GPT-2 are correlated with the toxic subspace, with stronger alignment in later layers and with more samples