IC-1166MPLUG-Owl's VQA accuracy on VQA-X increases from 68.30% to 74.48% when prompted with progressively higher-quality rationales generated by RAPPER
Kai-Po Chang, Chi-Pin Huang, Wei-Yuan Cheng, Fu-En Yang, Chien-Yi Wang, Yung-Hsuan Lai, Yu-Chiang Frank Wang
The paper tests MPLUG-Owl on the VQA-X test set under three input conditions: no rationale, a rationale from RAPPER's knowledge-distillation stage only, and a rationale from RAPPER's full two-stage pipeline (KD + RLNF). MPLUG-Owl's VQA accuracy rises monotonically across these conditions, from 68.30% with no rationale to 69.84% with KD-only rationales to 74.48% with full KD+RLNF rationales. This shows that MPLUG-Owl can leverage externally generated textual rationales to improve its visual question answering, and that the quality of the rationale (as shaped by the RLNF stage) matters.
The rationales are generated by the authors' own RAPPER pipeline, so the result is entangled with RAPPER's quality; it is not a general statement about all rationales. Only one VQA dataset (VQA-X) is tested.