Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jiarui, Hao, Yifan, Zhang, Hanning, Dong, Hanze, Xiong, Wei, Jiang, Nan, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
Scalable Chain of Thoughts via Elastic Reasoning
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
Fractured Chain-of-Thought Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
by: Chen, Qiguang, et al.
Published: (2026)
by: Chen, Qiguang, et al.
Published: (2026)
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation
by: Qiu, Yuli, et al.
Published: (2024)
by: Qiu, Yuli, et al.
Published: (2024)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)
by: Yao, Yao, et al.
Published: (2023)
VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
by: Gong, Xuan, et al.
Published: (2025)
by: Gong, Xuan, et al.
Published: (2025)
Understanding Before Reasoning: Enhancing Chain-of-Thought with Iterative Summarization Pre-Prompting
by: Zhu, Dong-Hai, et al.
Published: (2025)
by: Zhu, Dong-Hai, et al.
Published: (2025)
Faithful Logical Reasoning via Symbolic Chain-of-Thought
by: Xu, Jundong, et al.
Published: (2024)
by: Xu, Jundong, et al.
Published: (2024)
Efficient Reasoning via Chain of Unconscious Thought
by: Gong, Ruihan, et al.
Published: (2025)
by: Gong, Ruihan, et al.
Published: (2025)
Cognitive Loop of Thought: Reversible Hierarchical Markov Chain for Efficient Mathematical Reasoning
by: Zhang, Jia-Chen, et al.
Published: (2026)
by: Zhang, Jia-Chen, et al.
Published: (2026)
Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
by: Wang, Tianduo, et al.
Published: (2024)
by: Wang, Tianduo, et al.
Published: (2024)
FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean4
by: Yao, Jiarui, et al.
Published: (2025)
by: Yao, Jiarui, et al.
Published: (2025)
Statistical Rejection Sampling Improves Preference Optimization
by: Liu, Tianqi, et al.
Published: (2023)
by: Liu, Tianqi, et al.
Published: (2023)
Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
by: Deng, Jie, et al.
Published: (2026)
by: Deng, Jie, et al.
Published: (2026)
To Reason or Not to: Selective Chain-of-Thought in Medical Question Answering
by: Zhan, Zaifu, et al.
Published: (2026)
by: Zhan, Zaifu, et al.
Published: (2026)
Entropy-Regularized Process Reward Model
by: Zhang, Hanning, et al.
Published: (2024)
by: Zhang, Hanning, et al.
Published: (2024)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
by: Yu, Sheldon, et al.
Published: (2025)
by: Yu, Sheldon, et al.
Published: (2025)
Demystifying Long Chain-of-Thought Reasoning in LLMs
by: Yeo, Edward, et al.
Published: (2025)
by: Yeo, Edward, et al.
Published: (2025)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
by: Ma, Junyu, et al.
Published: (2025)
by: Ma, Junyu, et al.
Published: (2025)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
by: Zhang, Hongbo, et al.
Published: (2025)
by: Zhang, Hongbo, et al.
Published: (2025)
Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
by: Xiong, Juming, et al.
Published: (2026)
by: Xiong, Juming, et al.
Published: (2026)
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
by: Liu, Dong, et al.
Published: (2026)
by: Liu, Dong, et al.
Published: (2026)
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
by: Wu, Yuanhao, et al.
Published: (2025)
by: Wu, Yuanhao, et al.
Published: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
Faster Sampling via Stochastic Gradient Proximal Sampler
by: Huang, Xunpeng, et al.
Published: (2024)
by: Huang, Xunpeng, et al.
Published: (2024)
Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling
by: Zheng, Guangmin, et al.
Published: (2024)
by: Zheng, Guangmin, et al.
Published: (2024)
Short Chains, Deep Thoughts: Balancing Reasoning Efficiency and Intra-Segment Capability via Split-Merge Optimization
by: Gui, Runquan, et al.
Published: (2026)
by: Gui, Runquan, et al.
Published: (2026)
S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs
by: Du, Yanrui, et al.
Published: (2026)
by: Du, Yanrui, et al.
Published: (2026)
Enhancing Chain of Thought Prompting in Large Language Models via Reasoning Patterns
by: Zhang, Yufeng, et al.
Published: (2024)
by: Zhang, Yufeng, et al.
Published: (2024)
Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
by: Chen, Qiguang, et al.
Published: (2024)
by: Chen, Qiguang, et al.
Published: (2024)
Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
by: Alhazmi, Elaf, et al.
Published: (2026)
by: Alhazmi, Elaf, et al.
Published: (2026)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
by: Ye, Jiacheng, et al.
Published: (2024)
by: Ye, Jiacheng, et al.
Published: (2024)
Zero-Shot Chain-of-Thought Reasoning Guided by Evolutionary Algorithms in Large Language Models
by: Jin, Feihu, et al.
Published: (2024)
by: Jin, Feihu, et al.
Published: (2024)
Similar Items
-
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
by: Xiong, Wei, et al.
Published: (2025) -
Scalable Chain of Thoughts via Elastic Reasoning
by: Xu, Yuhui, et al.
Published: (2025) -
Fractured Chain-of-Thought Reasoning
by: Liao, Baohao, et al.
Published: (2025) -
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
by: Xiong, Wei, et al.
Published: (2025) -
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
by: Pang, Bo, et al.
Published: (2025)