IC-122Concept representations in Llama-2-7B, Gemma-7B, and Llama-2-13B become more consistent in deeper layers

Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani, Fan Yang, Mengnan Du

SourceBeyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution

The paper measures cosine similarity of concept vectors (derived via linear probing classifiers) at every layer of three LLMs. Across all three models, the average cosine similarity of concept vectors shows a general upward trend from early to penultimate layers, indicating that concept representations are less defined in early layers and become sharper and more consistent in deeper layers. Sampled vectors from the GCS distribution show the highest and most stable similarity across all layers. The pattern is remarkably consistent across the different model sizes and architectures.

Evidence
correlational
Key metric
cosine similarity among observed concept vectors ranging from approximately 0.8 to 0.9; sampled vectors within 1σ with cosine similarity larger than 0.93; observed-sampled similarity typically ranging from 0.88 to 0.93
Model
Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b, Llama-2-13B-Chat, Gemma Gemma-7B
Concepts
Depth-dependent structure
Methods
Linear Probing / Ridge regression linear probing / Linear probe / Linear probe fine-tuning / Linear regression probing / Linear ridge regression probes / Supervised probing / ERM linear probe [primary]
Related work
Zou et al. 2023 (Representation Engineering) [builds-on]
Related findings
IC-123
Extraction
automatic-extraction