SourceLinear Representations of Political Perspective Emerge in Large Language Models
To evaluate the political slant of 1,134 generated essays, the authors recruit 10 human annotators (3 Democrats, 4 Independents, 3 Republicans) from CloudResearch Survey to rate a random sample of 21 essays on a 7-point scale. GPT-4o is then prompted to rate the same essays on the identical scale. The intraclass correlation between GPT-4o's ratings and the averaged human ratings is 0.91, and the Spearman correlation is 0.952. The authors use this agreement as validation to deploy GPT-4o for rating the full set of essays.