SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yuyang, Cheng, Yi, Ying, Haochao, Du, Zhuoyun, Hu, Renjun, Shi, Xing, Lin, Wei, Wu, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models Could Be Rote Learners
by: Xu, Yuyang, et al.
Published: (2025)
by: Xu, Yuyang, et al.
Published: (2025)
Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
by: Cheng, Yi, et al.
Published: (2024)
by: Cheng, Yi, et al.
Published: (2024)
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
SSPO: Subsentence-level Policy Optimization
by: Yang, Kun, et al.
Published: (2025)
by: Yang, Kun, et al.
Published: (2025)
LLMs Can Simulate Standardized Patients via Agent Coevolution
by: Du, Zhuoyun, et al.
Published: (2024)
by: Du, Zhuoyun, et al.
Published: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
by: Lai, Xin, et al.
Published: (2024)
by: Lai, Xin, et al.
Published: (2024)
PoCo: A Self-Supervised Approach via Polar Transformation Based Progressive Contrastive Learning for Ophthalmic Disease Diagnosis
by: Wang, Jinhong, et al.
Published: (2024)
by: Wang, Jinhong, et al.
Published: (2024)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
by: Fei, Wu, et al.
Published: (2025)
by: Fei, Wu, et al.
Published: (2025)
Enabling Agents to Communicate Entirely in Latent Space
by: Du, Zhuoyun, et al.
Published: (2025)
by: Du, Zhuoyun, et al.
Published: (2025)
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
Versatile and Risk-Sensitive Cardiac Diagnosis via Graph-Based ECG Signal Representation
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
by: Hu, Renjun, et al.
Published: (2025)
by: Hu, Renjun, et al.
Published: (2025)
HDMoE: A Hierarchical Decoupling-Fusion Mixture-of-Experts Framework for Multimodal Cancer Survival Prediction
by: Wang, Huayi, et al.
Published: (2026)
by: Wang, Huayi, et al.
Published: (2026)
SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
Decouple, Reorganize, and Fuse: A Multimodal Framework for Cancer Survival Prediction
by: Wang, Huayi, et al.
Published: (2025)
by: Wang, Huayi, et al.
Published: (2025)
Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
by: Wang, Boxuan, et al.
Published: (2025)
by: Wang, Boxuan, et al.
Published: (2025)
Step-wise Rubric Rewards for LLM Reasoning
by: Xie, Weichu, et al.
Published: (2026)
by: Xie, Weichu, et al.
Published: (2026)
LegalReasoner: Step-wised Verification-Correction for Legal Judgment Reasoning
by: Shi, Weijie, et al.
Published: (2025)
by: Shi, Weijie, et al.
Published: (2025)
Step-level Value Preference Optimization for Mathematical Reasoning
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework
by: Huang, Kerui, et al.
Published: (2025)
by: Huang, Kerui, et al.
Published: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Self-Supervised One-Step Diffusion Refinement for Snapshot Compressive Imaging
by: Huang, Shaoguang, et al.
Published: (2024)
by: Huang, Shaoguang, et al.
Published: (2024)
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Process Supervision-Guided Policy Optimization for Code Generation
by: Dai, Ning, et al.
Published: (2024)
by: Dai, Ning, et al.
Published: (2024)
Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization
by: He, Shan, et al.
Published: (2026)
by: He, Shan, et al.
Published: (2026)
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
by: Shi, Liangtao, et al.
Published: (2025)
by: Shi, Liangtao, et al.
Published: (2025)
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
by: Wang, Tianduo, et al.
Published: (2024)
by: Wang, Tianduo, et al.
Published: (2024)
Self-Supervised Scalable Deep Compressed Sensing
by: Chen, Bin, et al.
Published: (2023)
by: Chen, Bin, et al.
Published: (2023)
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
by: Hu, Suhang, et al.
Published: (2025)
by: Hu, Suhang, et al.
Published: (2025)
SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
by: Peng, Liang, et al.
Published: (2025)
by: Peng, Liang, et al.
Published: (2025)
Structure-based RNA Design by Step-wise Optimization of Latent Diffusion Model
by: Si, Qi, et al.
Published: (2026)
by: Si, Qi, et al.
Published: (2026)
COAD: Constant-Time Planning for Continuous Goal Manipulation with Compressed Library and Online Adaptation
by: Shiyas, Adil, et al.
Published: (2026)
by: Shiyas, Adil, et al.
Published: (2026)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
Evaluating Step-by-Step Reasoning through Symbolic Verification
by: Zhang, Yi-Fan, et al.
Published: (2022)
by: Zhang, Yi-Fan, et al.
Published: (2022)
SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
by: Tang, Xiaqiang, et al.
Published: (2025)
by: Tang, Xiaqiang, et al.
Published: (2025)
StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
by: Zhou, Chenyu, et al.
Published: (2025)
by: Zhou, Chenyu, et al.
Published: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
TSO: Self-Training with Scaled Preference Optimization
by: Chen, Kaihui, et al.
Published: (2024)
by: Chen, Kaihui, et al.
Published: (2024)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Self-Supervised Learning in Deep Networks: A Pathway to Robust Few-Shot Classification
by: Xiao, Yuyang
Published: (2024)
by: Xiao, Yuyang
Published: (2024)
Similar Items
-
Large Language Models Could Be Rote Learners
by: Xu, Yuyang, et al.
Published: (2025) -
Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
by: Cheng, Yi, et al.
Published: (2024) -
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025) -
SSPO: Subsentence-level Policy Optimization
by: Yang, Kun, et al.
Published: (2025) -
LLMs Can Simulate Standardized Patients via Agent Coevolution
by: Du, Zhuoyun, et al.
Published: (2024)