anchor
Findings
- IC-078GPT-4's self-verification loop causes performance collapse due to high false negative rates in binary verification [compared-to]
- IC-1007LLMs cannot reliably self-verify or self-correct their own outputs without external tool feedback [compared-to]
- IC-1159GPT-4 is a strong inductive hypothesis proposer, achieving high accuracy on inductive reasoning benchmarks when its generated rules are applied by a symbolic interpreter [compared-to]
- IC-1599Self-repair at equivalent compute budget provides only modest and inconsistent gains over i.i.d. sampling for CodeLlama-13B-Instruct, GPT-3.5, and GPT-4 on HumanEval and APPS [context]
- IC-522LLMs perform correct example inference without inducing the correct rule, and this gap is robust to prompting methods, fact count, and scenario form [compared-to]
- IC-852Intrinsic self-correction without external feedback consistently degrades reasoning accuracy across GPT-3.5-turbo, GPT-4, GPT-4-turbo, and LLaMA-2-70B-chat [compared-to]
- IC-854The apparent self-correction improvement in constrained generation (Madaan et al., 2023) is an artefact of a sub-optimal initial prompt, not a genuine model capability [builds-on]
- IC-854The apparent self-correction improvement in constrained generation (Madaan et al., 2023) is an artefact of a sub-optimal initial prompt, not a genuine model capability [primary]