IC-1309GPT-4-0613 exhibits reduced output diversity relative to GPT-3.5: it outperforms on best@1 but underperforms on best@8 under CoT prompting
Alexander G Shypula, Aman Madaan, Yimeng Zeng, Uri Alon, Jacob R. Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, Amir Yazdanbakhsh
Under chain-of-thought prompting, GPT-4-0613 achieves 26.99% optimization rate at best@1 versus 21.37% for GPT-3.5, but at best@8 GPT-4 reaches only 42.74% opt / 1.58x speedup while GPT-3.5 reaches 43.05% opt / 1.60x. The authors attribute this to a lack of output diversity in GPT-4 despite using the same sampling hyper-parameters (temperature 0.7). The gap is small in absolute terms but consistent across the CoT condition.
The authors use the hedged language 'this may demonstrate a lack of output diversity'; the best@8 gap is small (0.31% opt, 0.02x speedup) and could reflect sampling variance rather than a systematic diversity difference.