IC-333Llama-3-405B overexplains math solutions with structured markdown headings and conversational tone compared to GPT-4o's concise formal notation, achieving 97% model-matching accuracy

Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez

SourceVibeCheck: Discover and Quantify Qualitative Differences in Large Language Models

On the MATH dataset with chain-of-thought prompting, VibeCheck compares GPT-4o and Llama-3-405B on questions where both answered correctly. Llama-3-405B organizes responses under markdown headings (e.g., '## Step 1'), adopts a conversational tone, and includes overly detailed step-by-step explanations with repetition. GPT-4o uses flowing narrative, formal tone, and frequent LaTeX/MathML notation. The top vibes (explanation detail sep 0.90, structural formatting sep 0.70, conciseness sep 0.51) achieve 97.09% model-matching accuracy and 72.79% preference-prediction accuracy. GPT-4o is favored in 76% of conversations.

Evidence
correlational
Key metric
Model matching accuracy 97.09%; preference prediction accuracy 72.79%; GPT-4o win rate 76%; separability scores: explanation and step-by-step detail 0.90, structural formatting 0.70, conciseness 0.51, efficiency of steps 0.42, mathematical notation use 0.33
Caveat
VibeCheck was run only on questions where both models answered correctly to reduce variance from incorrect examples. Only 5 vibes were found because the reduction step identified only 5 distinct vibes that could almost perfectly separate model outputs.
Model
GPT-4o, Llama 3 Llama-3-405B
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [supporting]
Related findings
IC-332, IC-334
Extraction
automatic-extraction