Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Monitoring Latent World States in Language Models with Propositional Probes
2025-01-22
· ICLR 2025 Spotlight ·
anchor
Findings
IC-044
Tulu-2-13B's internal activations contain a linearly decodable, faithful representation of input-context propositions that persists under prompt injection and backdoor attacks where outputs become unfaithful
IC-045
A 50-dimensional Hessian-identified subspace in Tulu-2-13B causally mediates entity-attribute binding, generalizing to three-entity contexts
IC-046
Tulu-2-13B exhibits gender bias in both its internal binding representation and its outputs, with the output-level bias being stronger than the representation-level bias
IC-047
Tulu-2-13B's entity-attribute binding partially relies on token order as a shortcut, degrading in nested orderings where order and semantic binding conflict