Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Databricks Dolly
anchor
Findings
IC-560
Llama3-8B-Instruct reliably distinguishes its own outputs from human outputs in self-recognition tasks, while Llama3-8B base performs at chance, indicating the ability is acquired during post-training.
[eval]
IC-563
The self-recognition vector's activation in Llama3-8B-Instruct is organized across depth: early layers (4-6) show diffuse perceptual activation to self-written text (present in both chat and base models), while layers 14-16 show a sharp decision-related peak at the output token that is present only in the chat model with role tags.
[eval]