Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CNN/DailyMail
anchor
Findings
IC-560
Llama3-8B-Instruct reliably distinguishes its own outputs from human outputs in self-recognition tasks, while Llama3-8B base performs at chance, indicating the ability is acquired during post-training.
[eval]
IC-563
The self-recognition vector's activation in Llama3-8B-Instruct is organized across depth: early layers (4-6) show diffuse perceptual activation to self-written text (present in both chat and base models), while layers 14-16 show a sharp decision-related peak at the output token that is present only in the chat model with role tags.
[eval]
IC-748
Pruned LLMs at ≥50% sparsity remain robust in-context retrievers and summarizers, with Vicuna-7B matching up to ~40% sparsity and Vicuna-13B up to ~50% sparsity in open-book settings
[eval]
IC-849
GPT-2 and T5-base exhibit vanishing expected gradients under RFT for inputs with small reward standard deviation, prevalent in 3 of 7 GRUE datasets, causing RFT to underperform SFT
[eval]