Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient Reasoning
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-698
GPT-3.5-turbo's CoT reasoning errors are correlated across different demonstration sets, while PoT errors are less correlated
IC-699
Llama2-13b produces significantly less consistent answers than GPT-3.5-turbo on complex reasoning tasks, making it unsuitable as a weaker LLM in a cascade
IC-700
GPT-4's reasoning accuracy degrades when provided with incorrect hints from a weaker model