Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Fictional Knowledge Dataset / Perez et al. (2022b) self-knowledge dataset
anchor
Findings
IC-1240
GPT-4-0314 achieves 85% zero-shot accuracy on situational-awareness questions about its own architecture and training
[eval]
IC-370
Knowledge entropy (sparsity of FFN memory coefficients) decreases consistently during pretraining for OLMo 1B, 7B, and Pythia 1.4B, and this decrease strongly correlates with reduced knowledge acquisition and increased forgetting in continual learning
[eval]
IC-371
Artificially resuscitating inactive memory vectors by scaling the up-projection matrix K improves knowledge acquisition and reduces forgetting, with the effect more pronounced for later-stage OLMo models
[eval]