When parametric memory and counter-memory are presented in different orders, the memorization ratio changes significantly for most models. PaLM2 varies from 38.6 (parametric memory first) to 72.2 (counter-memory first), a 33.6 percentage point swing. LLaMA2-7B varies from 33.3 to 82.8, a 49.5 percentage point swing. ChatGPT tends to favor the first evidence, while PaLM2 and LLaMA2-7B lean toward later evidence. GPT-4 shows minimal order sensitivity.
Evidence
correlational
Key metric
PaLM2: 38.6 (parametric first) vs 72.2 (counter first); LLaMA2-7B: 33.3 vs 82.8; ChatGPT: 46.7 vs 40.1; GPT-4: 60.9 vs 62.7
Caveat
Only four models (ChatGPT, GPT-4, PaLM2, LLaMA2-7B) were tested for order sensitivity; the order of evidence was the only variable changed.