Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
WikiText-2
anchor
Findings
IC-454
Most hidden trajectories in trained LLMs exhibit exponential growth in norm as a function of depth, a property that emerges with training
[eval]
IC-462
GPT-2 encodes toxicity in a low-dimensional linear subspace of its MLP layers, concentrated in higher layers
[source]
IC-463
DPO's first-step gradients in GPT-2 are correlated with the toxic subspace, with stronger alignment in later layers and with more samples
[source]
IC-528
The knowledge localization assumption fails for a large fraction of facts in GPT-2, Llama2-7B, and Llama3-8B, with 77% of facts classified as inconsistent knowledge in Llama3-8B
[eval]