The paper categorizes multifaceted and structural factual questions by the number of hops (1, 2, 3, or multi-hop) needed to verify them, sampling 1,490 data pieces from each dataset. For GPT-3.5-turbo under few-shot CoT, macro F1 drops steadily: from 51.3 at 1-hop to 45.2 at multi-hop for multifaceted questions, and from 40.6 at 1-hop to 30.2 at 3-hop for structural questions. The authors attribute this to the extended reasoning chain requiring heightened logical reasoning capabilities.