Logit lens

anchor

Decode intermediate activations through the model's output head to read what the model would predict at that depth.

Findings