Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Jie, Tong, Hanshuang, Li, Jun, Liang, Shining, Wu, Ning, Li, Hongzhi, Xie, Yutao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915773699588096
author Deng, Jie
Tong, Hanshuang
Li, Jun
Liang, Shining
Wu, Ning
Li, Hongzhi
Xie, Yutao
author_facet Deng, Jie
Tong, Hanshuang
Li, Jun
Liang, Shining
Wu, Ning
Li, Hongzhi
Xie, Yutao
contents Large language models (LLMs) have made impressive strides in mathematical reasoning, often fine-tuned using rejection sampling that retains only correct reasoning trajectories. While effective, this paradigm treats supervision as a binary filter that systematically excludes teacher-generated errors, leaving a gap in how reasoning failures are modeled during training. In this paper, we propose TrajFusion, a fine-tuning strategy that reframes rejection sampling as a structured supervision construction process. Specifically, TrajFusion forms fused trajectories that explicitly model trial-and-error reasoning by interleaving selected incorrect trajectories with reflection prompts and correct trajectories. The length of each fused sample is adaptively controlled based on the frequency and diversity of teacher errors, providing richer supervision for challenging problems while safely reducing to vanilla rejection sampling fine-tuning (RFT) when error signals are uninformative. TrajFusion requires no changes to the architecture or training objective. Extensive experiments across multiple math benchmarks demonstrate that TrajFusion consistently outperforms RFT, particularly on challenging and long-form reasoning problems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04391
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
Deng, Jie
Tong, Hanshuang
Li, Jun
Liang, Shining
Wu, Ning
Li, Hongzhi
Xie, Yutao
Computation and Language
Large language models (LLMs) have made impressive strides in mathematical reasoning, often fine-tuned using rejection sampling that retains only correct reasoning trajectories. While effective, this paradigm treats supervision as a binary filter that systematically excludes teacher-generated errors, leaving a gap in how reasoning failures are modeled during training. In this paper, we propose TrajFusion, a fine-tuning strategy that reframes rejection sampling as a structured supervision construction process. Specifically, TrajFusion forms fused trajectories that explicitly model trial-and-error reasoning by interleaving selected incorrect trajectories with reflection prompts and correct trajectories. The length of each fused sample is adaptively controlled based on the frequency and diversity of teacher errors, providing richer supervision for challenging problems while safely reducing to vanilla rejection sampling fine-tuning (RFT) when error signals are uninformative. TrajFusion requires no changes to the architecture or training objective. Extensive experiments across multiple math benchmarks demonstrate that TrajFusion consistently outperforms RFT, particularly on challenging and long-form reasoning problems.
title Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
topic Computation and Language
url https://arxiv.org/abs/2602.04391