Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
F1 score
Findings
IC-053
In GPT-2 small's l10h0 name mover queries, the io attribute is encoded with higher-magnitude features than the s attribute, and both are causally relevant, but SAEs preferentially learn io features due to the magnitude asymmetry
[eval]
IC-074
Released LLMs achieve F1 plan scores between 42.7 and 86.7 on the T-Eval plan task
[eval]