Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

2025-01-22 · ICLR 2025 Oral · anchor

Findings