Abstract:The ability to generate diverse solutions to a given problem is a hallmark of human creativity. This divergent reasoning is also crucial for machines, enhancing their robustness and enabling them to assist humans in many applications such as scientific discovery. However, existing approaches to multi-step reasoning with large language models (LLMs) have mostly focused only on reasoning accuracy, without further discovering more diverse valid solutions. For example, supervised fine-tuning can improve LLM reasoning quality, but requires extensive supervised data to capture the full range of possible solutions. Reinforcement learning aims to find limited highest-reward solutions while neglecting the solution diversity. To fill this gap, we propose Flow of Reasoning (FoR), an efficient diversity-seeking LLM finetuning method aimed at improving reasoning quality and diversity with minimal data. FoR formulates multi-step LLM reasoning as a Markovian flow on a DAG-structured reasoning graph. This formulation allows us to incorporate and adapt principled GFlowNet approaches, for finetuning LLMs to sample diverse reasoning paths with probabilities proportional to the (unnormalized) reward of target problems. Extensive experiments show that, with limited training examples (e.g., 15 examples), FoR enables the discovery of diverse, creative, high-quality solutions, greatly outperforming a wide range of existing inference and training methods across five challenging puzzle-solving tasks, including BlocksWorld (embodied reasoning), Game24 (math puzzle solving), Rubik's Cube (spatial reasoning), 1D-ARC (abstraction reasoning), and PrOntoQA (logical reasoning). Code is available at <a class="link-external link-https" href="https://github.com/Yu-Fangxu/FoR" rel="external noopener nofollow">this https URL</a>.

OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning

ReFT: Reasoning with Reinforced Fine-Tuning

OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

RAFT: Adapting Language Model to Domain Specific RAG

On the Impact of Fine-Tuning on Chain-of-Thought Reasoning

ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback

ALERT: Adapting Language Models to Reasoning Tasks

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

ReFT: Representation Finetuning for Language Models

Technical Report: Enhancing LLM Reasoning with Reward-guided Tree Search

Intuitive Fine-Tuning: Towards Unifying SFT and RLHF into a Single Process

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Flow of Reasoning:Training LLMs for Divergent Problem Solving with Minimal Examples

OPT-R: Exploring the Role of Explanations in Finetuning and Prompting for Reasoning Skills of Large Language Models

RLHF Workflow: From Reward Modeling to Online RLHF

Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards

AnyTaskTune: Advanced Domain-Specific Solutions through Task-Fine-Tuning

Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

A Comparative Study on Reasoning Patterns of OpenAI's o1 Model