Sparse autoencoder / Sparse autoencoders / K-sparse autoencoder / Topk sparse autoencoder / Scaling and Evaluating Sparse Autoencoders / Cunningham et al. 2023 (sparse autoencoders)

Train an overcomplete autoencoder with a sparsity penalty on model activations, so that individual learned features stand for interpretable directions.

Findings