Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

2025-01-22 · ICLR 2025 Poster · anchor

Findings