Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Neuron based Personality Trait Induction in Large Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-200
Personality-related neurons in Llama-3-8B-Instruct are concentrated in the deeper layers of the network
IC-201
Activating neuroticism-positive neurons in Llama-3-8B-Instruct causes the largest decline in general capabilities, while activating conscientiousness-positive neurons improves all benchmarks