The paper evaluates nine CLIP variants ranging from RN50 to ViT-L/14@336px on the STEM test set. Average accuracy ranges from 51.3% (ViT-B/16) to 54.9% (RN50x64), a spread of only 3.6 percentage points despite orders-of-magnitude differences in parameter count. The authors conclude that simply scaling up CLIP is insufficient and that new algorithmic advancements are needed.