The paper reports MMBench scores for five released vision-language models. InstructBLIP achieves the highest overall score (44.0), followed by LLaVA (38.7), VisualGLM (38.1), and OpenFlamingo (4.6). InstructBLIP leads in attribute reasoning (54.2) and coarse perception (56.4). OpenFlamingo scores very low across all sub-abilities, with 0.0 on relation reasoning. The numbers are sourced from the MMBench paper (Liu et al. 2023b).