IC-1147GPT-4 outperforms GPT-3.5 on all KITAB metrics but the gap is modest, with all-correctness below 35% for both, suggesting scale alone does not resolve constraint satisfaction
Marah I Abdin, Suriya Gunasekar, Varun Chandrasekaran, Jerry Li, Mert Yuksekgonul, Rahee Ghosh Peshawaria, Ranjita Naik, Besmira Nushi
Across all conditions and constraint types, GPT-4 achieves higher satisfaction, completeness, and all-correctness than GPT-3.5, but the differences are not dramatic. In the best condition (with-context), GPT-4 achieves 0.08 all-correctness versus 0.07 for GPT-3.5. The paper explicitly states that 'the difference between the two llms is not as dramatic showing that scale alone may not address the filtering with constraints problem.' This holds across popularity bins and constraint types.
Evidence
correlational
Key metric
All-correctness with-context: GPT-4 0.08, GPT-3.5 0.07; all-correctness no-context: GPT-4 0.08, GPT-3.5 0.07; all correctness remains notably lower than 35% across all conditions
Caveat
Only two models from the same family (OpenAI) are compared; the claim about 'scale' is limited to this specific model family and does not generalize to cross-architecture comparisons.