Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-939
CLIP reward landscapes are well-shaped for photorealistic environments but poorly shaped for abstract renderings
IC-940
CLIP can specify 5 of 8 complex humanoid tasks from single-sentence prompts, failing on tasks requiring discrimination of subtle body-pose differences
IC-941
CLIP reward model quality scales with model size, with a sharp phase transition between ViT-H/14 and ViT-BigG/14 for the humanoid kneeling task