IC-940CLIP can specify 5 of 8 complex humanoid tasks from single-sentence prompts, failing on tasks requiring discrimination of subtle body-pose differences

Juan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez, David Lindner

SourceVision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Using CLIP ViT-BigG/14 as a zero-shot reward model, the authors train a MuJoCo humanoid on 8 tasks specified by single-sentence text prompts. Five tasks (kneeling, lotus position, standing up, arms raised, doing splits) achieve 100% human-evaluated success. Three tasks fail: hands on hips (64%), standing on one leg (0%), and arms crossed (0%). The authors hypothesize the failures stem from CLIP's inability to discriminate between visually similar body poses (e.g., hands on hips vs. standing) and from physical limitations of the MuJoCo simulation (round feet making one-leg standing difficult). Goal-baseline regularization does not rescue the failing tasks.

Evidence
correlational
Key metric
kneeling 100%, lotus position 100%, standing up 100%, arms raised 100%, doing splits 100%, hands on hips 64%, standing on one leg 0%, arms crossed 0% (human-evaluated success rate over 100 trajectories, 4 random seeds)
Caveat
Human evaluation by a single author; the authors note their protocol is 'very basic and might be biased.' The standing-on-one-leg failure may be due to MuJoCo physics (round feet) rather than CLIP's visual discrimination.
Model
CLIP / CLIP-ViT (LC)
Concepts
Failure mode
Datasets
MuJoCo Humanoid-v4 [eval]
Methods
EPIC distance [eval]
Related work
Du et al. 2023 [compared-to]
Related findings
IC-939, IC-941
Extraction
automatic-extraction