Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Big-Bench Hard
anchor
Note
no anchor recorded: the citing paper attaches arXiv 2206.04615, which is BIG-bench ("Beyond the Imitation Game") and not BIG-Bench Hard. Verified in the bibliography of corpus/text/st77ShxP1K.txt, which prints no identifier for BBH itself
Findings
IC-029
Large language models show conformity to group answers in multi-agent interactions
[source]
IC-1342
Assigning socio-demographic personas to LLMs causes significant reasoning performance degradation across all four models studied, manifesting as both explicit abstentions and implicit reasoning errors
[eval]
IC-698
GPT-3.5-turbo's CoT reasoning errors are correlated across different demonstration sets, while PoT errors are less correlated
[eval]
IC-699
Llama2-13b produces significantly less consistent answers than GPT-3.5-turbo on complex reasoning tasks, making it unsuitable as a weaker LLM in a cascade
[eval]
IC-700
GPT-4's reasoning accuracy degrades when provided with incorrect hints from a weaker model
[eval]