IC-807Language-conditioned robot policies RT-1 and RT-2 fail to generalize to unseen manipulation tasks, achieving only 16.7% and 11.1% success rates respectively across 7 novel skills
Jiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu, Montserrat Gonzalez Arenas, Kanishka Rao, Wenhao Yu, Chuyuan Fu, Keerthana Gopalakrishnan, Zhuo Xu, Priya Sundaresan, Peng Xu, Hao Su, Karol Hausman, Chelsea Finn, Quan Vuong, Ted Xiao
The paper evaluates RT-1 and RT-2 on 7 unseen manipulation skills (place fruit, upright and move, move within drawer, restock drawer, pick from chair, fold towel, swivel chair) that require combining seen motions in novel ways or generalizing to wholly unseen motions. Both language-conditioned policies perform poorly: RT-1 reaches 16.7% overall success and RT-2 reaches 11.1%, across 64 total trials. The authors note that language conditioning is insufficient because the new tasks have semantically unseen language instructions, even when the underlying arm motions were partially seen during training. RT-2, despite being trained on internet-scale VQA data, does not outperform the smaller RT-1 on these novel-motion tasks.
The evaluation uses human-drawn trajectory sketches for the RT-Trajectory baseline but language instructions for RT-1 and RT-2, so the comparison is between different conditioning modalities rather than a controlled ablation of conditioning type on the same model.