IC-561A linear direction in the residual stream at layer 16 of Llama3-8B-Instruct is causally necessary and sufficient for self-authorship claims: steering with it achieves 100% control over authorship assertions, and projecting it out reduces claims by 50-60%.

Christopher Ackerman, Nina Panickssery

SourceInspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct

Using the contrastive pairs method, the authors isolate a unit vector from the mean difference in residual stream activations between self- and other-written texts. Steering with this vector at multipliers 3-6 on layers 14-16 causes the model to claim authorship of any text (including human-written) at 100% rate, or deny authorship of any text at 100% rate. Projecting the vector out of the residual stream at layer 16 during generation reduces self-authorship claims by 50-60% across three datasets. The effect is specific: steering on a dummy named-entity-recognition task has no effect at layer 16, ruling out a generic agreement mechanism.

Evidence
interventional
Key metric
Steering: 100% effectiveness at multipliers 3-6, layers 14-16. Positive steering multiplier 10, layer 16: model claims authorship of non-self text ~80% of the time vs ~35% unsteered. Ablation: zeroing out vector at layer 16 reduces self-authorship claims from ~50% to under 30% (50-60% decrease) across three datasets.
Caveat
The vector was derived from 290 contrastive pairs, a relatively small sample. The authors note that high multipliers in early layers produce degenerate output. The vector's effect on a dummy task is not entirely absent (modest negative-direction effects at high multipliers), which the authors attribute to overlap with negation concepts.
Model
Llama 3 8B Instruct
Concepts
Linear representation, Depth-dependent structure
Datasets
SAD [eval]
Methods
Contrastive Pairs Method [primary], Activation steering / Mean steering / PCA steering [primary], Tuned Lens [validation]
Related work
Turner et al. 2024 (Activation Addition) [builds-on], Zou et al. 2023 (Representation Engineering) [builds-on]
Related findings
IC-560, IC-562, IC-563
Extraction
automatic-extraction