IC-609The last transformer layer of Mistral 7B, Llama 3.2 1B, and Llama 3.2 3B shows anomalous trajectory statistics inconsistent with the linear drift-plus-noise pattern of intermediate layers

Raphaël Sarfati, Toni J.B. Liu, Nicolas Boulle, Christopher Earls

SourceLines of Thought in Large Language Models

When the same trajectory analysis is applied to Mistral 7B (32 layers), Llama 3.2 1B (16 layers), and Llama 3.2 3B (28 layers), the linear drift-plus-Gaussian-noise model holds for all intermediate layers but breaks down at the final layer. In Mistral 7B, the last layer (32) is misaligned such that linear extrapolation produces an error much larger than expected. In Llama 3.2 1B and 3B, the last layer shows out-of-distribution mean and variance of the residuals δx. The authors conjecture this may be an effect of re-alignment or fine-tuning, since the first and last layers are most exposed to perturbations that do not propagate deep into the stack. All models also show minor deviations in the very first layers.

Evidence
observational
Caveat
The reason for the anomaly is not identified; the authors state it is 'not immediately evident' and 'perhaps worth investigating further.' The effect is described qualitatively (larger variance, misalignment) without a single quantitative threshold.
Model
Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1, Llama 3.2
Concepts
Depth-dependent structure
Related findings
IC-608
Extraction
automatic-extraction