IC-170GPT-4, GPT-3.5, and Claude-3.5-Sonnet rely heavily on parametric knowledge in RAG settings, producing ungrounded responses with high answered ratios and low trust-scores
Maojia Song, Shang Hong Sim, Rishabh Bhardwaj, Hai Leong Chieu, Navonil Majumder, Soujanya Poria
The paper measures how much released closed-source models rely on their internal parametric knowledge versus the provided documents in a RAG setting. GPT-4 answers 86.81% of ASQA questions and 73.40% of QAMPARI questions, with parametric knowledge scores (sparam) of 12.71 and 13.05. Claude-3.5-Sonnet shows a similar pattern with 84.60% AR% on ASQA and sparam 12.99. This reliance on parametric knowledge rather than grounding in documents results in lower trust-scores (GPT-4: 69.72 ASQA, 40.35 QAMPARI, 49.33 ELI5; Claude-3.5: 64.36, 39.78, 25.92) despite higher answer correctness scores, because the responses are not supported by the cited documents.
sparam only partially captures parametric knowledge utilization; it does not account for cases where the document contains the answer and the model still relies on parametric knowledge to generate the correct answer