The paper tests CLIP's vulnerability to typographical attacks by overlaying the word 'ocean' in a white box on 100 forest-class images from UC Merced Land Use. In a zero-shot 5-class setting (ocean, forest, runway, parking, residential), the unmodified CLIP correctly classifies 94% of clean forest images but 99% of the attacked images are misclassified as ocean. On ImageNet (10 categories, 50 images each), the attack reduces accuracy from 99.80% to 54.00%. Randomly removing the same number of tokens recovers only 17% of attacked images, while removing tokens that ViTE interprets as text-related recovers 97%.