IC-014Sparse autoencoders can identify 'uncertainty' directions in the residual stream before an answer, which are predictive of incorrect responses.

Javier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, Neel Nanda

SourceDo I Know This Entity? Knowledge Awareness and Hallucinations in Language Models

The authors identify SAE latents in the residual stream at the end-of-instruction token that differentiate between correct and incorrect responses. These directions, found in Gemma 2 2b and Gemma 2 9b, appear to signal model uncertainty. The highest-scoring latent in Gemma 2 2b achieves a 73.2% AUROC and an F1 score of 72% in classifying correct vs. incorrect responses on a test set. An analysis of the latent's activations on a large text corpus shows it fires on text related to uncertainty or undisclosed information.

Evidence
correlational
Key metric
Top latent for Gemma 2 2b: 73.2% AUROC, F1 score 72%. The t-statistic is used to select top latents from a training set.
Caveat
This analysis excludes instances where the model refuses to answer. The relationship is correlational; the study does not intervene on these directions to demonstrate causality. Results are based on a specific set of entity types.
Model
Gemma 2 2B, 9B
Concepts
Linear representation
Datasets
Wikidata / WikidataRecent [source], FineWeb [source]
Methods
Sparse autoencoder / Sparse autoencoders / K-sparse autoencoder / Topk sparse autoencoder / Scaling and Evaluating Sparse Autoencoders / Cunningham et al. 2023 (sparse autoencoders) [primary]
Extraction
automatic-extraction