Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-332
Llama-3-70B exhibits a friendlier, funnier, and less ethics-focused style than GPT-4 and Claude-3-Opus on Chatbot Arena, and these vibes predict model identity at 80% and user preference at 59% accuracy
IC-333
Llama-3-405B overexplains math solutions with structured markdown headings and conversational tone compared to GPT-4o's concise formal notation, achieving 97% model-matching accuracy
IC-334
GPT-4V produces more poetic, emotion-focused image captions compared to Gemini-1.5-Flash's literal descriptions, with 99% model-matching accuracy on COCO