IC-540Spatial understanding is the most challenging dimension for code-based models while numerical understanding is most challenging for direct image models
Breaking down correctness by understanding type reveals a model-type-specific weakness. For code-based models (GPT-4o, Llama, AutoTikZ), spatial understanding scores are the lowest: GPT-4o drops from ~4.0 on attribute binding to 3.35-3.47 on spatial. For direct image models (Stable Diffusion, DALL-E), numerical understanding is the weakest: both score below 1.8 on numerical versus ~2.7 on attribute. This discrepancy means combined tasks involving the weak dimension (e.g., numerical & spatial) receive the lowest scores across all models.