SourceResidual Stream Analysis with Multi-Layer SAEs
The authors relax the assumption that the residual stream basis is the same at every layer by applying pre-trained tuned-lens transformations (Belrose et al., 2023) to the activation vectors before passing them to the MLSAE encoder. Contrary to the expectation that aligning the basis would increase multi-layer activity, the tuned-lens approach decreases the fraction of total variance explained by individual latents, keeping it approximately constant between 37% and 41% as the expansion factor increases. Conversely, the single-token variance ratio increases, meaning the single-prompt heatmaps become more spread out. This indicates that part of the observed multi-layer structure in the standard setting is attributable to the basis change across layers rather than to genuinely shared features.