Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Linear Representations of Political Perspective Emerge in Large Language Models
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-501
Linear probes on middle-layer attention heads of Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 predict US lawmakers' DW-Nominate ideology scores with Spearman correlations around 0.85
IC-502
Linear probes trained on US lawmaker ideology generalize to predict Ad Fontes media slant scores when the same models simulate news outlets
IC-503
Adding probe regression coefficients to attention head activations steers Llama-2-7B-Chat, Mistral-7B-Instruct-v0.1, and Vicuna-7B-v1.5 toward more liberal or conservative generated text
IC-504
GPT-4o (gpt-4o-2024-08-06) rates the political slant of LLM-generated essays in close agreement with politically balanced human annotators