Abstract:The ability to generate diverse solutions to a given problem is a hallmark of human creativity. This divergent reasoning is also crucial for machines, enhancing their robustness and enabling them to assist humans in many applications such as scientific discovery. However, existing approaches to multi-step reasoning with large language models (LLMs) have mostly focused only on reasoning accuracy, without further discovering more diverse valid solutions. For example, supervised fine-tuning can improve LLM reasoning quality, but requires extensive supervised data to capture the full range of possible solutions. Reinforcement learning aims to find limited highest-reward solutions while neglecting the solution diversity. To fill this gap, we propose Flow of Reasoning (FoR), an efficient diversity-seeking LLM finetuning method aimed at improving reasoning quality and diversity with minimal data. FoR formulates multi-step LLM reasoning as a Markovian flow on a DAG-structured reasoning graph. This formulation allows us to incorporate and adapt principled GFlowNet approaches, for finetuning LLMs to sample diverse reasoning paths with probabilities proportional to the (unnormalized) reward of target problems. Extensive experiments show that, with limited training examples (e.g., 15 examples), FoR enables the discovery of diverse, creative, high-quality solutions, greatly outperforming a wide range of existing inference and training methods across five challenging puzzle-solving tasks, including BlocksWorld (embodied reasoning), Game24 (math puzzle solving), Rubik's Cube (spatial reasoning), 1D-ARC (abstraction reasoning), and PrOntoQA (logical reasoning). Code is available at <a class="link-external link-https" href="https://github.com/Yu-Fangxu/FoR" rel="external noopener nofollow">this https URL</a>.

Flow of Reasoning: Efficient Training of LLM Policy with Divergent Thinking

Flow of Reasoning:Training LLMs for Divergent Problem Solving with Minimal Examples

Concise and Organized Perception Facilitates Large Language Models for Deductive Reasoning.

Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning

GFlowNet Fine-tuning for Diverse Correct Solutions in Mathematical Reasoning Tasks

HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows

Proof Flow: Preliminary Study on Generative Flow Network Language Model Tuning for Formal Reasoning

Unleashing the Creative Mind: Language Model As Hierarchical Policy For Improved Exploration on Challenging Problem Solving

Can LLMs Reason in the Wild with Programs?

Concise and Organized Perception Facilitates Reasoning in Large Language Models

Diversity of Thought Improves Reasoning Abilities of LLMs

P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains

Democratizing Reasoning Ability: Tailored Learning from Large Language Model

Enhancing Logical Reasoning in Large Language Models to Facilitate Legal Applications

Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning

LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models

ReFT: Reasoning with Reinforced Fine-Tuning

Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph

Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Automatic Curriculum Expert Iteration for Reliable LLM Reasoning