IC-527A 16 million latent sparse autoencoder substituted into GPT-4 yields a language modeling loss corresponding to 10% of GPT-4's pretraining compute
Leo Gao, Tom Dupre la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, Jeffrey Wu
The authors train a 16 million latent topk autoencoder on GPT-4's residual stream activations (layer 5/6) for 40 billion tokens. When the autoencoder reconstruction replaces the original residual stream during the forward pass, the resulting language modeling loss corresponds to a model trained with 10% of GPT-4's pretraining compute. This indicates that a sparse set of 16M features captures a substantial fraction of GPT-4's next-token prediction behavior, though the zero-ablation fidelity of the same autoencoder is 98.2%.
Evidence
interventional
Key metric
"language modeling loss corresponding to 10% of the pretraining compute of gpt-4"; zero-ablation fidelity 98.2%
Caveat
The 16M autoencoder was not trained to the l(n) convergence token budget due to compute constraints, so the reconstruction may not be optimal. The context length is 64 tokens. The 10% figure is a relative comparison to GPT-4's own pretraining compute, not an absolute loss value.