Light Dark Wikidata / WikidataRecent anchor · artifact
Note a live knowledge base rather than a fixed release, so what a study drew from it is defined by the study and not by a version; the artifact is the base itself and the anchor is the paper describing it Findings IC-012 Sparse autoencoders uncover entity recognition directions in Gemma 2 and Llama 3.1 models that are causally relevant for knowledge refusal. [source] IC-013 Entity recognition directions regulate attention to entity tokens in attribute extraction heads in Gemma and Llama models. [source] IC-014 Sparse autoencoders can identify 'uncertainty' directions in the residual stream before an answer, which are predictive of incorrect responses. [source] IC-1232 GPT-2 XL and GPT-J exhibit knowledge conflict when subjected to reverse and composite knowledge edits, with ROME and MEMIT showing near-total failure on reverse edits [source] IC-1233 GPT-2 XL and GPT-J exhibit irreversible knowledge distortion after round-editing, with the effect being more severe when the edit target is semantically distant from the true labels [source] IC-1261 LLaMA-2 attention to constraint tokens correlates with factual correctness, and a linear probe on these attention weights predicts factual errors comparably to model confidence [source] IC-1262 LLaMA-2 factual query accuracy improves with entity popularity and decreases with query constrainedness, with larger models showing better performance on less popular and more constrained queries [source] IC-1263 LLaMA-2 7B and 13B attention signal for predicting factual errors is available by approximately 50% of layers, enabling early stopping without performance degradation, while LLaMA-2 70B shows a slight performance drop [source] IC-1553 GPT-J, GPT-2-XL, and Llama-13B decode approximately 48% of tested relations via a linear transformation on the subject representation, and this structure causally influences predictions [source] IC-409 Knowledge editing methods correct verified hallucinations in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B far less effectively than their scores on existing benchmarks suggest [eval] IC-409 Knowledge editing methods correct verified hallucinations in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B far less effectively than their scores on existing benchmarks suggest [source]