IC-443Model inconsistency on probing questions negatively correlates with skill-slice accuracy (r = -0.675), with models contradicting themselves more often on skills where they perform poorly

Mazda Moayeri, Vidhisha Balachandran, Varun Chandrasekaran, Safoora Yousefi, Thomas FEL, Soheil Feizi, Besmira Nushi, Neel Joshi, Vibhav Vineet

SourceUnearthing Skill-level Insights for Understanding Trade-offs of Foundation Models

The paper generates 20 probing questions per skill targeting a single skill in isolation, obtains 5 responses per probe per model, and measures the rate of self-contradiction. Across three sets of 20 skills (lowest, median, highest slice accuracy), models contradict themselves at significantly higher rates for low-accuracy skills. The overall Pearson correlation between probe inconsistency rate and skill-slice accuracy is r = -0.675. Skills where all models contradict themselves more often than not include musical note identification, eye direction analysis, and enumerating objects.

Evidence
correlational
Key metric
r = -0.675 correlation between probe inconsistency rate and skill-slice accuracy
Caveat
Probe inconsistency rate can only be a factor of 0.05 since the implementation checks 20 claims per probed skill, limiting the resolution of the correlation.
Model
GPT-4o, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Claude 3.5 Sonnet
Concepts
Failure mode
Datasets
MMLU Pro [eval], MMMU [eval], MathVista [eval], MMC [eval], MMVP [eval], MMBench [eval], MMT-Bench [eval], MME [eval], MM-VET [eval], Seed-Bench / SeedBench [eval], Vibe-Eval [eval]
Related work
Wang et al. 2023 (self-consistency) [builds-on]
Related findings
IC-440, IC-441, IC-442
Extraction
automatic-extraction