Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Hypothesis Refinement / Iterative hypothesis refinement
anchor
Findings
IC-1157
GPT-4 and other LMs show a large gap between rule induction and rule application, with task accuracy dropping to near zero on MiniScan when the LM itself applies its own proposed rules
[primary]
IC-1158
GPT-4 and other LMs are brittle to noisy exemplars and unfamiliar output representations, with performance degrading sharply even under minimal perturbation
[primary]
IC-1159
GPT-4 is a strong inductive hypothesis proposer, achieving high accuracy on inductive reasoning benchmarks when its generated rules are applied by a symbolic interpreter
[primary]
IC-522
LLMs perform correct example inference without inducing the correct rule, and this gap is robust to prompting methods, fact count, and scenario form
[compared-to]