SourceTypedThinker: Diversify Large Language Model Reasoning with Typed Thinking
The paper prompts Mistral 7B Instruct with each of four reasoning types (deductive, inductive, abductive, analogical) on four benchmarks and samples up to 10 solutions per type at temperature 1. A problem is counted as exclusively solvable by a type if at least one solution under that type is correct but no solution under any other type is. Figure 1 shows that on every benchmark, a non-zero percentage of problems fall into this category for at least one non-deductive type, meaning the model cannot solve them regardless of how many times it is sampled with the wrong reasoning type. The authors note that the problems solvable by different types do not completely overlap, and that repeated sampling with high temperature does not overcome the limitation when the wrong type is used.