IC-915In Pythia-70M and Pythia-160M, individual MLP hidden neurons are activated by multiple irrelevant token combinations (pattern superposition)

Yuandong Tian, Yiping Wang, Zhenyu Zhang, Beidi Chen, Simon Shaolei Du

SourceJoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and Attention

The paper demonstrates that a single neuron in the MLP hidden layers of Pythia-70M and Pythia-160M can be activated by multiple unrelated token combinations. For example, the same neuron fires for both 'every morning' and 'in the realm of physics', which are semantically unrelated phrases. This 'pattern superposition' arises because the neuron's weight vector can encode multiple orthogonal query-pattern pairs simultaneously, as predicted by the JOMA framework's implicit bias mechanism.

Evidence
observational
Caveat
Demonstrated with a small number of illustrative examples (Figure 10) rather than a systematic quantitative measurement across all neurons.
Model
Pythia
Related findings
IC-913, IC-914
Extraction
automatic-extraction