Abstract:As Large Language Models (LLMs) scale up and gain powerful Chain-of-Thoughts (CoTs) reasoning abilities, practical resource constraints drive efforts to distill these capabilities into more compact Smaller Language Models (SLMs). We find that CoTs consist mainly of simple reasoning forms, with a small proportion ($\approx 4.7\%$) of key reasoning steps that truly impact conclusions. However, previous distillation methods typically involve supervised fine-tuning student SLMs only on correct CoTs data produced by teacher LLMs, resulting in students struggling to learn the key reasoning steps, instead imitating the teacher's reasoning forms and making errors or omissions on these steps. To address these issues, drawing an analogy to human learning, where analyzing mistakes according to correct solutions often reveals the crucial steps leading to successes or failures, we propose mistak\textbf{E}-\textbf{D}riven key reason\textbf{I}ng step distilla\textbf{T}ion (\textbf{EDIT}), a novel method that further aids SLMs learning key reasoning steps rather than mere simple fine-tuning. Firstly, to expose these crucial steps in CoTs, we design specific prompts to generate dual CoTs data with similar reasoning paths but divergent conclusions. Then, we apply the minimum edit distance algorithm on the dual CoTs data to locate these key steps and optimize the likelihood of these steps. Extensive experiments validate the effectiveness of EDIT across both in-domain and out-of-domain benchmark reasoning datasets. Further analysis shows that EDIT can generate high-quality CoTs with more correct key reasoning steps. Notably, we also explore how different mistake patterns affect performance and find that EDIT benefits more from logical errors than from knowledge or mathematical calculation errors in dual CoTs\footnote{Code can be found at \url{<a class="link-external link-https" href="https://github.com/C-W-D/EDIT" rel="external noopener nofollow">this https URL</a>}}.

PaD: Program-aided Distillation Can Teach Small Models Reasoning Better Than Chain-of-thought Fine-tuning

Concise and Organized Perception Facilitates Large Language Models for Deductive Reasoning.

Mixed Distillation Helps Smaller Language Model Better Reasoning

Mixed Distillation Helps Smaller Language Models Reason Better

Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model

Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation

Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation

Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step

Teaching Small Language Models Reasoning Through Counterfactual Distillation

Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation

Effective Distillation of Table-based Reasoning Ability from LLMs

Distilling Mathematical Reasoning Capabilities into Small Language Models

Distilling Algorithmic Reasoning from LLMs via Explaining Solution Programs

SIKeD: Self-guided Iterative Knowledge Distillation for mathematical reasoning

Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks

Mentor-KD: Making Small Language Models Better Multi-step Reasoners

Logic Distillation: Learning from Code Function by Function for Planning and Decision-making

MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models

TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-Reflection

DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language Models

Enhancing Code Generation Performance of Smaller Models by Distilling the Reasoning Ability of LLMs