The paper measures cosine similarity of concept vectors (derived via linear probing classifiers) at every layer of three LLMs. Across all three models, the average cosine similarity of concept vectors shows a general upward trend from early to penultimate layers, indicating that concept representations are less defined in early layers and become sharper and more consistent in deeper layers. Sampled vectors from the GCS distribution show the highest and most stable similarity across all layers. The pattern is remarkably consistent across the different model sizes and architectures.
Evidence
correlational
Key metric
cosine similarity among observed concept vectors ranging from approximately 0.8 to 0.9; sampled vectors within 1σ with cosine similarity larger than 0.93; observed-sampled similarity typically ranging from 0.88 to 0.93