Abstract:Complex multi-step reasoning tasks, such as solving mathematical problems or generating code, remain a significant hurdle for even the most advanced large language models (LLMs). Verifying LLM outputs with an Outcome Reward Model (ORM) is a standard inference-time technique aimed at enhancing the reasoning performance of LLMs. However, this still proves insufficient for reasoning tasks with a lengthy or multi-hop reasoning chain, where the intermediate outcomes are neither properly rewarded nor penalized. Process supervision addresses this limitation by assigning intermediate rewards during the reasoning process. To date, the methods used to collect process supervision data have relied on either human annotation or per-step Monte Carlo estimation, both prohibitively expensive to scale, thus hindering the broad application of this technique. In response to this challenge, we propose a novel divide-and-conquer style Monte Carlo Tree Search (MCTS) algorithm named \textit{OmegaPRM} for the efficient collection of high-quality process supervision data. This algorithm swiftly identifies the first error in the Chain of Thought (CoT) with binary search and balances the positive and negative examples, thereby ensuring both efficiency and quality. As a result, we are able to collect over 1.5 million process supervision annotations to train a Process Reward Model (PRM). Utilizing this fully automated process supervision alongside the weighted self-consistency algorithm, we have enhanced the instruction tuned Gemini Pro model's math reasoning performance, achieving a 69.4\% success rate on the MATH benchmark, a 36\% relative improvement from the 51\% base model performance. Additionally, the entire process operates without any human intervention, making our method both financially and computationally cost-effective compared to existing methods.

Exploring Equation As a Better Intermediate Meaning Representation for Numerical Reasoning of Large Language Models

Exploring Equation as a Better Intermediate Meaning Representation for Numerical Reasoning

Specialized Mathematical Solving by a Step-By-Step Expression Chain Generation

Reflection of Thought: Inversely Eliciting Numerical Reasoning in Language Models via Solving Linear Systems

Using Intermediate Representations to Solve Math Word Problems.

Enhancing Numerical Reasoning with the Guidance of Reliable Reasoning Processes

When Do Program-of-Thought Works for Reasoning?

Solving Math Word Problems by Combining Language Models With Symbolic Solvers

Improving Arithmetic Reasoning Ability of Large Language Models through Relation Tuples, Verification and Dynamic Feedback

An Expression Tree Decoding Strategy for Mathematical Equation Generation

Improve Mathematical Reasoning in Language Models by Automated Process Supervision

LLM-SR: Scientific Equation Discovery via Programming with Large Language Models

Reasoning in Large Language Models Through Symbolic Math Word Problems

Distilling Algorithmic Reasoning from LLMs via Explaining Solution Programs

Large language models for automatic equation discovery of nonlinear dynamics

LLM4ED: Large Language Models for Automatic Equation Discovery

Targeted training for numerical reasoning with large language models

Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models

Small Language Models are Equation Reasoners

Interpreting and Improving Large Language Models in Arithmetic Calculation