anchor
Findings
- IC-1167GPT-4 Code Interpreter's mathematical reasoning accuracy is positively correlated with code usage frequency, with its iterative code generation and self-debugging mechanism as the primary driver of its 69.69% zero-shot MATH accuracy [eval]
- IC-1168For GPT-4 Code Interpreter, code-based self-verification improves MATH accuracy to 73.54% while natural language self-verification slightly degrades it to 69.29% relative to the 69.69% base [eval]
- IC-1169CodeLlama-7B and CodeLlama-34B show improvement with CSV prompting on GSM8K and MATH but at much lower absolute accuracy than GPT-4 Code Interpreter [eval]
- IC-1170GPT-3.5, Llama2, PaLM2, and GPT-4 are susceptible to a CoT-prompting backdoor attack (BadChain) on complex reasoning tasks, with stronger reasoning models showing higher attack success rates [eval]
- IC-1330All 20 evaluated LLMs improve in multi-turn task-solving with additional tool-use turns and GPT-4-simulated language feedback [eval]
- IC-1523Five AI assistants (Claude-1.3, Claude-2.0, GPT-3.5-turbo, GPT-4, Llama-2-70B-Chat) consistently exhibit sycophancy across four varied free-form text-generation tasks [eval]
- IC-798GPT-4 achieves 61.6% on MATH with tool-integrated reasoning prompting, exceeding PaL (51.8%) and CoT (42.5%) [eval]