IC-502Linear probes trained on US lawmaker ideology generalize to predict Ad Fontes media slant scores when the same models simulate news outlets

Junsol Kim, James Evans, Aaron Schein

SourceLinear Representations of Political Perspective Emerge in Large Language Models

Using the same linear probes trained to predict DW-Nominate scores from lawmaker-prompt activations, the authors evaluate whether they can predict the Ad Fontes political slant of 400 US news outlets when the models are prompted to generate statements from those outlets. No new probes are trained; the existing DW-Nominate probes are applied directly to the news-outlet activations. The ensembled predictions (top 32 heads) achieve Spearman correlations of 0.798 (Llama-2-7B-Chat), 0.764 (Mistral-7B-Instruct), and 0.720 (Vicuna-7B), demonstrating that the representation captures a generalizable liberal–conservative axis rather than memorized entity-specific scores.

Evidence
correlational
Key metric
Spearman ρ_cv (k=32 ensemble): 0.798 (Llama-2-7B-Chat), 0.764 (Mistral-7B-Instruct), 0.720 (Vicuna-7B)
Caveat
The generalization is to US media outlets only; cross-national party ideology (411 parties) yields a much lower ρ=0.531 for Llama-2-7B-Chat, indicating the axis is US-specific.
Model
Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct-v0.1, Vicuna Vicuna-7B-v1.5
Concepts
Linear representation
Datasets
Ad Fontes Media [eval]
Methods
Linear Probing / Ridge regression linear probing / Linear probe / Linear probe fine-tuning / Linear regression probing / Linear ridge regression probes / Supervised probing / ERM linear probe [primary]
Related findings
IC-501, IC-503, IC-504
Extraction
automatic-extraction