SourceNot All Language Model Features Are One-Dimensionally Linear
While GPT-2-small contains circular representations of days, months, and years in its layer-7 activations (discovered via SAE clustering), the model scores only 8/49 on the weekdays task and 10/144 on the months task, which the authors describe as trivial accuracy. This contrasts sharply with Mistral 7B (31/49, 125/144) and Llama 3 8B (29/49, 143/144) on the same tasks, suggesting that the mere presence of a circular representation is insufficient for the model to use it in computation.