Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Huimin, Mao, Xin, Li, Feng-Lin, Wu, Xiaobao, Chen, Wang, Zhang, Wei, Luu, Anh Tuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation
por: Xu, Huimin, et al.
Publicado: (2025)
por: Xu, Huimin, et al.
Publicado: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
por: Lai, Xin, et al.
Publicado: (2024)
por: Lai, Xin, et al.
Publicado: (2024)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
por: Mao, Xin, et al.
Publicado: (2024)
por: Mao, Xin, et al.
Publicado: (2024)
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
por: Srey, Ponhvoan, et al.
Publicado: (2025)
por: Srey, Ponhvoan, et al.
Publicado: (2025)
Step-level Value Preference Optimization for Mathematical Reasoning
por: Chen, Guoxin, et al.
Publicado: (2024)
por: Chen, Guoxin, et al.
Publicado: (2024)
As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback Loss
por: Mao, Xin, et al.
Publicado: (2024)
por: Mao, Xin, et al.
Publicado: (2024)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
por: Xu, Yuyang, et al.
Publicado: (2025)
por: Xu, Yuyang, et al.
Publicado: (2025)
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
por: Lu, Zimu, et al.
Publicado: (2024)
por: Lu, Zimu, et al.
Publicado: (2024)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
por: Fei, Wu, et al.
Publicado: (2025)
por: Fei, Wu, et al.
Publicado: (2025)
A Survey on Neural Topic Models: Methods, Applications, and Challenges
por: Wu, Xiaobao, et al.
Publicado: (2024)
por: Wu, Xiaobao, et al.
Publicado: (2024)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
por: Yuan, Youliang, et al.
Publicado: (2025)
por: Yuan, Youliang, et al.
Publicado: (2025)
Towards the TopMost: A Topic Modeling System Toolkit
por: Wu, Xiaobao, et al.
Publicado: (2023)
por: Wu, Xiaobao, et al.
Publicado: (2023)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
por: Xu, Huimin, et al.
Publicado: (2026)
por: Xu, Huimin, et al.
Publicado: (2026)
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
por: Xu, Jundong, et al.
Publicado: (2025)
por: Xu, Jundong, et al.
Publicado: (2025)
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
por: Deng, Yihe, et al.
Publicado: (2025)
por: Deng, Yihe, et al.
Publicado: (2025)
Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme Detection
por: Pan, Fengjun, et al.
Publicado: (2025)
por: Pan, Fengjun, et al.
Publicado: (2025)
Are LLMs Good Zero-Shot Fallacy Classifiers?
por: Pan, Fengjun, et al.
Publicado: (2024)
por: Pan, Fengjun, et al.
Publicado: (2024)
Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
por: Feng, Yichao, et al.
Publicado: (2025)
por: Feng, Yichao, et al.
Publicado: (2025)
KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
por: Nguyen, Cong-Duy, et al.
Publicado: (2024)
por: Nguyen, Cong-Duy, et al.
Publicado: (2024)
GroupDPO: Memory efficient Group-wise Direct Preference Optimization
por: Leng, Jixuan, et al.
Publicado: (2026)
por: Leng, Jixuan, et al.
Publicado: (2026)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
por: Wu, Junkang, et al.
Publicado: (2024)
por: Wu, Junkang, et al.
Publicado: (2024)
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
por: Chen, Shaolong, et al.
Publicado: (2026)
por: Chen, Shaolong, et al.
Publicado: (2026)
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
por: Srey, Ponhvoan, et al.
Publicado: (2026)
por: Srey, Ponhvoan, et al.
Publicado: (2026)
AKEW: Assessing Knowledge Editing in the Wild
por: Wu, Xiaobao, et al.
Publicado: (2024)
por: Wu, Xiaobao, et al.
Publicado: (2024)
Learning Uncertainty from Sequential Internal Dispersion in Large Language Models
por: Srey, Ponhvoan, et al.
Publicado: (2026)
por: Srey, Ponhvoan, et al.
Publicado: (2026)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
por: Cao, Lang, et al.
Publicado: (2024)
por: Cao, Lang, et al.
Publicado: (2024)
Uncertainty-Aware Step-wise Verification with Generative Reward Models
por: Ye, Zihuiwen, et al.
Publicado: (2025)
por: Ye, Zihuiwen, et al.
Publicado: (2025)
Topic Modeling as Multi-Objective Contrastive Optimization
por: Nguyen, Thong, et al.
Publicado: (2024)
por: Nguyen, Thong, et al.
Publicado: (2024)
Modeling Dynamic Topics in Chain-Free Fashion by Evolution-Tracking Contrastive Learning and Unassociated Word Exclusion
por: Wu, Xiaobao, et al.
Publicado: (2024)
por: Wu, Xiaobao, et al.
Publicado: (2024)
Step-wise Rubric Rewards for LLM Reasoning
por: Xie, Weichu, et al.
Publicado: (2026)
por: Xie, Weichu, et al.
Publicado: (2026)
MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering
por: Siyue, Zhang, et al.
Publicado: (2024)
por: Siyue, Zhang, et al.
Publicado: (2024)
SOD: Step-wise On-policy Distillation for Small Language Model Agents
por: Zhong, Qiyong, et al.
Publicado: (2026)
por: Zhong, Qiyong, et al.
Publicado: (2026)
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
por: Wu, Xiaobao
Publicado: (2025)
por: Wu, Xiaobao
Publicado: (2025)
Vision-and-Language Pretraining
por: Nguyen, Thong, et al.
Publicado: (2022)
por: Nguyen, Thong, et al.
Publicado: (2022)
InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
por: Wu, Xiaobao, et al.
Publicado: (2023)
por: Wu, Xiaobao, et al.
Publicado: (2023)
MUR: Momentum Uncertainty guided Reasoning for Large Language Models
por: Yan, Hang, et al.
Publicado: (2025)
por: Yan, Hang, et al.
Publicado: (2025)
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths
por: Chia, Yew Ken, et al.
Publicado: (2024)
por: Chia, Yew Ken, et al.
Publicado: (2024)
ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations
por: Pham, Quang Hieu, et al.
Publicado: (2025)
por: Pham, Quang Hieu, et al.
Publicado: (2025)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
por: Yang, Xiliang, et al.
Publicado: (2025)
por: Yang, Xiliang, et al.
Publicado: (2025)
FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
por: Wu, Xiaobao, et al.
Publicado: (2024)
por: Wu, Xiaobao, et al.
Publicado: (2024)
Ejemplares similares
-
SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation
por: Xu, Huimin, et al.
Publicado: (2025) -
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
por: Lai, Xin, et al.
Publicado: (2024) -
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
por: Mao, Xin, et al.
Publicado: (2024) -
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
por: Srey, Ponhvoan, et al.
Publicado: (2025) -
Step-level Value Preference Optimization for Mathematical Reasoning
por: Chen, Guoxin, et al.
Publicado: (2024)