GPT-4 Code Interpreter for Math
First page

Paper summary
A zero-shot prompting technique for GPT-4 Code Interpreter that dramatically boosts math-reasoning accuracy via code self-verification.
Ask this paper
01
Code-as-verifier prompting: Explicitly encourages GPT-4 Code Interpreter to use code for self-verification of intermediate and final answers.
02
69.7% on MATH: Achieves 69.7% zero-shot accuracy on the MATH dataset - a 27.5-point improvement over vanilla GPT-4 (42.2%).
03
Execution-grounded reasoning: Code execution provides a high-fidelity verification signal that vanilla CoT lacks, reducing hallucinated intermediate steps.
04
Tool-use template: Establishes a template for tool-augmented reasoning that would generalize to many later math-LLM recipes.