IC-297Applying tuned-lens transformations to the residual stream decreases the apparent multi-layer SAE feature activity from 54–88% to 37–41% of total variance

Tim Lawson, Lucy Farnik, Conor Houghton, Laurence Aitchison

SourceResidual Stream Analysis with Multi-Layer SAEs

The authors relax the assumption that the residual stream basis is the same at every layer by applying pre-trained tuned-lens transformations (Belrose et al., 2023) to the activation vectors before passing them to the MLSAE encoder. Contrary to the expectation that aligning the basis would increase multi-layer activity, the tuned-lens approach decreases the fraction of total variance explained by individual latents, keeping it approximately constant between 37% and 41% as the expansion factor increases. Conversely, the single-token variance ratio increases, meaning the single-prompt heatmaps become more spread out. This indicates that part of the observed multi-layer structure in the standard setting is attributable to the basis change across layers rather than to genuinely shared features.

Evidence
correlational
Key metric
tuned-lens fraction of total variance explained by individual latents: between 37% and 41% (vs. 54–88% standard); ratio remains approximately constant as expansion factor increases
Caveat
Pre-trained tuned lenses were only available for Pythia-70m, Pythia-160m, and Pythia-410m; no tuned lens existed for Pythia-1b at the time of writing, so results are limited to three model sizes.
Model
Pythia
Concepts
Method artefact
Datasets
The Pile [source]
Methods
Tuned Lens [primary]
Related work
Eliciting Latent Predictions from Transformers with the Tuned Lens [builds-on]
Related findings
IC-296
Extraction
automatic-extraction