Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ScImage: How good are multimodal large language models at scientific text-to-image generation?
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-539
GPT-4o in text-code-image mode achieves the highest scores on SCIMAGE but remains below 4 on all three evaluation dimensions, and all models degrade substantially on prompts requiring combined understanding types
IC-540
Spatial understanding is the most challenging dimension for code-based models while numerical understanding is most challenging for direct image models
IC-541
Code-based output (Python or TikZ) produces images with notably higher scientific style scores than direct image generation from Stable Diffusion and DALL-E