Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Towards Understanding Factual Knowledge of Large Language Models
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-733
CoT prompting improves factual accuracy for instruction-tuned LLMs but degrades it for non-instruction-tuned LLMs such as OPT, BLOOM, and LLaMA
IC-734
GPT-3.5-turbo's factual verification F1 decreases as the number of reasoning hops required to validate a claim increases
IC-735
GPT-3.5-turbo's factual verification performance drops substantially under adversarial modifications, with man-made adversarial examples causing the largest decline
IC-736
Vicuna-13B outperforms Vicuna-7B on factual knowledge tasks by 5.4% on average