SourceNeuron based Personality Trait Induction in Large Language Models
The paper identifies approximately 20,000 personality-related neurons in Llama-3-8B-Instruct by measuring activation probability differences between opposing trait aspects. When the distribution of these neurons is plotted across the 29 transformer layers, they are primarily concentrated in the deeper layers (roughly layers 17–29), with far fewer in the early layers. The authors note this is consistent with prior work showing that FFNs in the last few layers are more crucial for conceptual knowledge, suggesting the model's understanding of personality-like concepts emerges progressively with depth.