Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Deng, Jie, Tong, Hanshuang, Li, Jun, Liang, Shining, Wu, Ning, Li, Hongzhi, Xie, Yutao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
por: Deng, Jie, et al.
Publicado: (2026)
por: Deng, Jie, et al.
Publicado: (2026)
PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
por: Yin, Shangjian, et al.
Publicado: (2025)
por: Yin, Shangjian, et al.
Publicado: (2025)
Ploutos: Towards interpretable stock movement prediction with financial large language model
por: Tong, Hanshuang, et al.
Publicado: (2024)
por: Tong, Hanshuang, et al.
Publicado: (2024)
Evaluating Mathematical Reasoning Beyond Accuracy
por: Xia, Shijie, et al.
Publicado: (2024)
por: Xia, Shijie, et al.
Publicado: (2024)
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
por: Chen, Nuo, et al.
Publicado: (2023)
por: Chen, Nuo, et al.
Publicado: (2023)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
por: Tong, Yuxuan, et al.
Publicado: (2024)
por: Tong, Yuxuan, et al.
Publicado: (2024)
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
por: Xiong, Wei, et al.
Publicado: (2025)
por: Xiong, Wei, et al.
Publicado: (2025)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
por: Yang, Shiping, et al.
Publicado: (2025)
por: Yang, Shiping, et al.
Publicado: (2025)
Reasons to Reject? Aligning Language Models with Judgments
por: Xu, Weiwen, et al.
Publicado: (2023)
por: Xu, Weiwen, et al.
Publicado: (2023)
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning
por: Yu, Fei, et al.
Publicado: (2025)
por: Yu, Fei, et al.
Publicado: (2025)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
por: Li, Junjie, et al.
Publicado: (2026)
por: Li, Junjie, et al.
Publicado: (2026)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
por: Yao, Jiarui, et al.
Publicado: (2025)
por: Yao, Jiarui, et al.
Publicado: (2025)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
por: Cao, Lang, et al.
Publicado: (2024)
por: Cao, Lang, et al.
Publicado: (2024)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
por: Liu, Weihao, et al.
Publicado: (2024)
por: Liu, Weihao, et al.
Publicado: (2024)
Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement
por: Zhang, Ying, et al.
Publicado: (2026)
por: Zhang, Ying, et al.
Publicado: (2026)
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
por: Zhang, Zhihan, et al.
Publicado: (2024)
por: Zhang, Zhihan, et al.
Publicado: (2024)
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
por: Rao, Jun, et al.
Publicado: (2025)
por: Rao, Jun, et al.
Publicado: (2025)
Toward Automated Robustness Evaluation of Mathematical Reasoning
por: Hou, Yutao, et al.
Publicado: (2025)
por: Hou, Yutao, et al.
Publicado: (2025)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
por: Huang, Hongzhi, et al.
Publicado: (2025)
por: Huang, Hongzhi, et al.
Publicado: (2025)
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
por: Ding, Bowen, et al.
Publicado: (2025)
por: Ding, Bowen, et al.
Publicado: (2025)
Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations
por: Chen, Nuo, et al.
Publicado: (2023)
por: Chen, Nuo, et al.
Publicado: (2023)
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
por: Pei, Qizhi, et al.
Publicado: (2025)
por: Pei, Qizhi, et al.
Publicado: (2025)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
por: Zhang, Di, et al.
Publicado: (2024)
por: Zhang, Di, et al.
Publicado: (2024)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
por: Wan, Guangya, et al.
Publicado: (2024)
por: Wan, Guangya, et al.
Publicado: (2024)
Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning
por: Zhang, Kepu, et al.
Publicado: (2025)
por: Zhang, Kepu, et al.
Publicado: (2025)
Statistical Rejection Sampling Improves Preference Optimization
por: Liu, Tianqi, et al.
Publicado: (2023)
por: Liu, Tianqi, et al.
Publicado: (2023)
Examining False Positives under Inference Scaling for Mathematical Reasoning
por: Wang, Yu, et al.
Publicado: (2025)
por: Wang, Yu, et al.
Publicado: (2025)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
por: Zhao, Yuze, et al.
Publicado: (2026)
por: Zhao, Yuze, et al.
Publicado: (2026)
Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
por: Hyun, Lee, et al.
Publicado: (2025)
por: Hyun, Lee, et al.
Publicado: (2025)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
por: Liu, Weihao, et al.
Publicado: (2025)
por: Liu, Weihao, et al.
Publicado: (2025)
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
por: Wu, Yuanhao, et al.
Publicado: (2025)
por: Wu, Yuanhao, et al.
Publicado: (2025)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
por: Zhang, Jiaqiao, et al.
Publicado: (2026)
por: Zhang, Jiaqiao, et al.
Publicado: (2026)
MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning
por: Li, Nianqi, et al.
Publicado: (2024)
por: Li, Nianqi, et al.
Publicado: (2024)
AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback
por: Gao, Zhitao, et al.
Publicado: (2026)
por: Gao, Zhitao, et al.
Publicado: (2026)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
por: Son, Guijin, et al.
Publicado: (2025)
por: Son, Guijin, et al.
Publicado: (2025)
Constrained Adaptive Rejection Sampling
por: Parys, Paweł, et al.
Publicado: (2025)
por: Parys, Paweł, et al.
Publicado: (2025)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
por: Zhao, Jun, et al.
Publicado: (2024)
por: Zhao, Jun, et al.
Publicado: (2024)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
por: Wang, Yiming, et al.
Publicado: (2024)
por: Wang, Yiming, et al.
Publicado: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
por: Li, Zhen, et al.
Publicado: (2025)
por: Li, Zhen, et al.
Publicado: (2025)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
por: Kuang, Jiayi, et al.
Publicado: (2025)
por: Kuang, Jiayi, et al.
Publicado: (2025)
Ejemplares similares
-
ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
por: Deng, Jie, et al.
Publicado: (2026) -
PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
por: Yin, Shangjian, et al.
Publicado: (2025) -
Ploutos: Towards interpretable stock movement prediction with financial large language model
por: Tong, Hanshuang, et al.
Publicado: (2024) -
Evaluating Mathematical Reasoning Beyond Accuracy
por: Xia, Shijie, et al.
Publicado: (2024) -
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
por: Chen, Nuo, et al.
Publicado: (2023)