IC-1166MPLUG-Owl's VQA accuracy on VQA-X increases from 68.30% to 74.48% when prompted with progressively higher-quality rationales generated by RAPPER

Kai-Po Chang, Chi-Pin Huang, Wei-Yuan Cheng, Fu-En Yang, Chien-Yi Wang, Yung-Hsuan Lai, Yu-Chiang Frank Wang

SourceRAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering

The paper tests MPLUG-Owl on the VQA-X test set under three input conditions: no rationale, a rationale from RAPPER's knowledge-distillation stage only, and a rationale from RAPPER's full two-stage pipeline (KD + RLNF). MPLUG-Owl's VQA accuracy rises monotonically across these conditions, from 68.30% with no rationale to 69.84% with KD-only rationales to 74.48% with full KD+RLNF rationales. This shows that MPLUG-Owl can leverage externally generated textual rationales to improve its visual question answering, and that the quality of the rationale (as shaped by the RLNF stage) matters.

Evidence
correlational
Key metric
VQA accuracy on VQA-X: 68.30% (x=none), 69.84% (x=r', KD only), 74.48% (x=r, KD+RLNF)
Caveat
The rationales are generated by the authors' own RAPPER pipeline, so the result is entangled with RAPPER's quality; it is not a general statement about all rationales. Only one VQA dataset (VQA-X) is tested.
Model
mPLUG-Owl
Datasets
VQA-X [eval]
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [primary]
Extraction
automatic-extraction