IC-539GPT-4o in text-code-image mode achieves the highest scores on SCIMAGE but remains below 4 on all three evaluation dimensions, and all models degrade substantially on prompts requiring combined understanding types
Seven models are evaluated on 404 scientific text-to-image prompts by 11 human scientists on correctness, relevance, and scientific style (1-5 scale). GPT-4o in both TikZ and Python code mode leads all models by more than 1.3 points on correctness, yet still scores below 4 on every dimension. All other models cluster between 1.7 and 2.2 on correctness. When prompts require all three understanding types (spatial, numerical, attribute) simultaneously, even GPT-4o records its lowest scores (3.13), and the average across all models drops to 2.21.
Evidence
correlational
Key metric
GPT-4o_tikz: correctness 3.50, relevance 3.67, scientific style 3.75; GPT-4o_python: correctness 3.51, relevance 3.40, scientific style 3.93; all other models correctness between 1.7 and 2.2; combined three-type average 2.21
Caveat
Compile errors are penalised with a score of 0, which substantially lowers code-model scores; Llama 3.1 8B has ~28% compilation error rate. Sample sizes for some object types are small (e.g., 4 for tables, 8 for matrices).