IC-540Spatial understanding is the most challenging dimension for code-based models while numerical understanding is most challenging for direct image models

Leixin Zhang, Steffen Eger, Yinjie Cheng, WEIHE ZHAI, Jonas Belouadi, Fahimeh Moafian, Zhixue Zhao

SourceScImage: How good are multimodal large language models at scientific text-to-image generation?

Breaking down correctness by understanding type reveals a model-type-specific weakness. For code-based models (GPT-4o, Llama, AutoTikZ), spatial understanding scores are the lowest: GPT-4o drops from ~4.0 on attribute binding to 3.35-3.47 on spatial. For direct image models (Stable Diffusion, DALL-E), numerical understanding is the weakest: both score below 1.8 on numerical versus ~2.7 on attribute. This discrepancy means combined tasks involving the weak dimension (e.g., numerical & spatial) receive the lowest scores across all models.

Evidence
correlational
Key metric
GPT-4o_tikz: attribute 4.11, numerical 3.49, spatial 3.35; Stable Diffusion: attribute 2.75, numerical 1.73, spatial 2.06; DALL-E: attribute 2.68, numerical 1.77, spatial 2.13
Caveat
Sample sizes per understanding type range from 40 to 80 prompts, limiting statistical power for individual cells.
Model
GPT-4o, Llama 3.1 8B, AutoTikZ / DataTikZ, Stable Diffusion, DALL-E
Concepts
Failure mode
Related findings
IC-539, IC-541
Extraction
automatic-extraction