IC-807Language-conditioned robot policies RT-1 and RT-2 fail to generalize to unseen manipulation tasks, achieving only 16.7% and 11.1% success rates respectively across 7 novel skills

Jiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu, Montserrat Gonzalez Arenas, Kanishka Rao, Wenhao Yu, Chuyuan Fu, Keerthana Gopalakrishnan, Zhuo Xu, Priya Sundaresan, Peng Xu, Hao Su, Karol Hausman, Chelsea Finn, Quan Vuong, Ted Xiao

SourceRT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

The paper evaluates RT-1 and RT-2 on 7 unseen manipulation skills (place fruit, upright and move, move within drawer, restock drawer, pick from chair, fold towel, swivel chair) that require combining seen motions in novel ways or generalizing to wholly unseen motions. Both language-conditioned policies perform poorly: RT-1 reaches 16.7% overall success and RT-2 reaches 11.1%, across 64 total trials. The authors note that language conditioning is insufficient because the new tasks have semantically unseen language instructions, even when the underlying arm motions were partially seen during training. RT-2, despite being trained on internet-scale VQA data, does not outperform the smaller RT-1 on these novel-motion tasks.

Evidence
correlational
Key metric
RT-1 overall success 16.7%, RT-2 overall success 11.1%, across 7 unseen skills and 64 trials (Table 4: RT-1 17%, RT-2 11%)
Caveat
The evaluation uses human-drawn trajectory sketches for the RT-Trajectory baseline but language instructions for RT-1 and RT-2, so the comparison is between different conditioning modalities rather than a controlled ablation of conditioning type on the same model.
Model
RT-1, RT-2
Concepts
Failure mode
Datasets
RT-1 Demonstration Dataset [train]
Methods
Behavior Cloning [supporting]
Related work
RT-1 [compared-to], RT-2 [compared-to]
Extraction
automatic-extraction