Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Jie, Tong, Hanshuang, Li, Jun, Liang, Shining, Wu, Ning, Li, Hongzhi, Xie, Yutao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
von: Deng, Jie, et al.
Veröffentlicht: (2026)
von: Deng, Jie, et al.
Veröffentlicht: (2026)
PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
von: Yin, Shangjian, et al.
Veröffentlicht: (2025)
von: Yin, Shangjian, et al.
Veröffentlicht: (2025)
Ploutos: Towards interpretable stock movement prediction with financial large language model
von: Tong, Hanshuang, et al.
Veröffentlicht: (2024)
von: Tong, Hanshuang, et al.
Veröffentlicht: (2024)
Evaluating Mathematical Reasoning Beyond Accuracy
von: Xia, Shijie, et al.
Veröffentlicht: (2024)
von: Xia, Shijie, et al.
Veröffentlicht: (2024)
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
von: Tong, Yuxuan, et al.
Veröffentlicht: (2024)
von: Tong, Yuxuan, et al.
Veröffentlicht: (2024)
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
Reasons to Reject? Aligning Language Models with Judgments
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2025)
von: Yu, Fei, et al.
Veröffentlicht: (2025)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
von: Li, Junjie, et al.
Veröffentlicht: (2026)
von: Li, Junjie, et al.
Veröffentlicht: (2026)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
von: Cao, Lang, et al.
Veröffentlicht: (2024)
von: Cao, Lang, et al.
Veröffentlicht: (2024)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement
von: Zhang, Ying, et al.
Veröffentlicht: (2026)
von: Zhang, Ying, et al.
Veröffentlicht: (2026)
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
von: Rao, Jun, et al.
Veröffentlicht: (2025)
von: Rao, Jun, et al.
Veröffentlicht: (2025)
Toward Automated Robustness Evaluation of Mathematical Reasoning
von: Hou, Yutao, et al.
Veröffentlicht: (2025)
von: Hou, Yutao, et al.
Veröffentlicht: (2025)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
von: Ding, Bowen, et al.
Veröffentlicht: (2025)
von: Ding, Bowen, et al.
Veröffentlicht: (2025)
Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
von: Zhang, Di, et al.
Veröffentlicht: (2024)
von: Zhang, Di, et al.
Veröffentlicht: (2024)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning
von: Zhang, Kepu, et al.
Veröffentlicht: (2025)
von: Zhang, Kepu, et al.
Veröffentlicht: (2025)
Statistical Rejection Sampling Improves Preference Optimization
von: Liu, Tianqi, et al.
Veröffentlicht: (2023)
von: Liu, Tianqi, et al.
Veröffentlicht: (2023)
Examining False Positives under Inference Scaling for Mathematical Reasoning
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
von: Zhao, Yuze, et al.
Veröffentlicht: (2026)
von: Zhao, Yuze, et al.
Veröffentlicht: (2026)
Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
von: Hyun, Lee, et al.
Veröffentlicht: (2025)
von: Hyun, Lee, et al.
Veröffentlicht: (2025)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
von: Wu, Yuanhao, et al.
Veröffentlicht: (2025)
von: Wu, Yuanhao, et al.
Veröffentlicht: (2025)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
von: Zhang, Jiaqiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaqiao, et al.
Veröffentlicht: (2026)
MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning
von: Li, Nianqi, et al.
Veröffentlicht: (2024)
von: Li, Nianqi, et al.
Veröffentlicht: (2024)
AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback
von: Gao, Zhitao, et al.
Veröffentlicht: (2026)
von: Gao, Zhitao, et al.
Veröffentlicht: (2026)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
von: Son, Guijin, et al.
Veröffentlicht: (2025)
von: Son, Guijin, et al.
Veröffentlicht: (2025)
Constrained Adaptive Rejection Sampling
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
von: Kuang, Jiayi, et al.
Veröffentlicht: (2025)
von: Kuang, Jiayi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
von: Deng, Jie, et al.
Veröffentlicht: (2026) -
PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
von: Yin, Shangjian, et al.
Veröffentlicht: (2025) -
Ploutos: Towards interpretable stock movement prediction with financial large language model
von: Tong, Hanshuang, et al.
Veröffentlicht: (2024) -
Evaluating Mathematical Reasoning Beyond Accuracy
von: Xia, Shijie, et al.
Veröffentlicht: (2024) -
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
von: Chen, Nuo, et al.
Veröffentlicht: (2023)