IC-332Llama-3-70B exhibits a friendlier, funnier, and less ethics-focused style than GPT-4 and Claude-3-Opus on Chatbot Arena, and these vibes predict model identity at 80% and user preference at 59% accuracy

Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez

SourceVibeCheck: Discover and Quantify Qualitative Differences in Large Language Models

Using pairwise battles from Chatbot Arena, VibeCheck discovers qualitative axes (vibes) along which Llama-3-70B differs from GPT-4 and Claude-3-Opus. The top vibes include Llama's use of typographic emphasis (sep 0.64), humor (sep 0.62), examples (sep 0.61), and lower ethical consideration (sep 0.53). Llama uses personal pronouns ('I', 'we', 'you') 3x more than GPT/Claude. A logistic regression over these vibes achieves 80.34% model-matching accuracy and 59.34% preference-prediction accuracy overall, with higher accuracy on writing prompts (77.19% MM, 62.04% PP) than STEM (68.71% MM, 57.31% PP).

Evidence
correlational
Key metric
Model matching accuracy 80.34% (overall), 68.71% (STEM), 77.19% (writing); preference prediction accuracy 59.34% (overall), 57.31% (STEM), 62.04% (writing); Cohen's kappa 0.46; separability scores: typographic emphasis 0.64, humor 0.62, examples 0.61, ethical consideration 0.53
Caveat
The paper notes that LLM judges (GPT-4o-mini, Llama-3-70B) are often incorrect in their predictions and have biases like favoring their own outputs. Running VibeCheck multiple times can lead to different vibes and results, making exact reproduction difficult.
Model
Llama 3 70B, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3 Opus
Datasets
Chatbot Arena [eval]
Methods
Logistic regression / Linear decoder (logistic regression) / Logistic classification / Logistic regression head [eval], Cohen's kappa [eval]
Related work
Chatbot Arena [context]
Related findings
IC-333, IC-334
Extraction
automatic-extraction