Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-012
Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal.
IC-013
Entity recognition directions regulate attention to entity tokens in attribute extraction heads in Gemma and Llama models.
IC-014
Sparse autoencoders can identify 'uncertainty' directions in the residual stream before an answer, which are predictive of incorrect responses.