Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Tianduo, Li, Shichen, Lu, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL
by: Liu, Hanbing, et al.
Published: (2025)
by: Liu, Hanbing, et al.
Published: (2025)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
by: Zhang, Hongbo, et al.
Published: (2025)
by: Zhang, Hongbo, et al.
Published: (2025)
Self-Harmonized Chain of Thought
by: Jin, Ziqi, et al.
Published: (2024)
by: Jin, Ziqi, et al.
Published: (2024)
TinyLlama: An Open-Source Small Language Model
by: Zhang, Peiyuan, et al.
Published: (2024)
by: Zhang, Peiyuan, et al.
Published: (2024)
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
by: Wang, Tianduo, et al.
Published: (2025)
by: Wang, Tianduo, et al.
Published: (2025)
Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
by: Chen, Xinghao, et al.
Published: (2025)
by: Chen, Xinghao, et al.
Published: (2025)
Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs
by: Nguyen, Minh-Vuong, et al.
Published: (2024)
by: Nguyen, Minh-Vuong, et al.
Published: (2024)
Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
by: Chen, Qiguang, et al.
Published: (2024)
by: Chen, Qiguang, et al.
Published: (2024)
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
by: Paul, Debjit, et al.
Published: (2024)
by: Paul, Debjit, et al.
Published: (2024)
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
by: Sun, Renliang, et al.
Published: (2025)
by: Sun, Renliang, et al.
Published: (2025)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Causal Sufficiency and Necessity Improves Chain-of-Thought Reasoning
by: Yu, Xiangning, et al.
Published: (2025)
by: Yu, Xiangning, et al.
Published: (2025)
Improving Chain-of-Thought Reasoning via Quasi-Symbolic Abstractions
by: Ranaldi, Leonardo, et al.
Published: (2025)
by: Ranaldi, Leonardo, et al.
Published: (2025)
Chain-of-Thought Reasoning Without Prompting
by: Wang, Xuezhi, et al.
Published: (2024)
by: Wang, Xuezhi, et al.
Published: (2024)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023)
by: Khalifa, Muhammad, et al.
Published: (2023)
Rethinking Chain-of-Thought from the Perspective of Self-Training
by: Wu, Zongqian, et al.
Published: (2024)
by: Wu, Zongqian, et al.
Published: (2024)
BPO: Revisiting Preference Modeling in Direct Preference Optimization
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)
by: Yao, Yao, et al.
Published: (2023)
Efficient Reasoning for LLMs through Speculative Chain-of-Thought
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
Efficient Reasoning via Chain of Unconscious Thought
by: Gong, Ruihan, et al.
Published: (2025)
by: Gong, Ruihan, et al.
Published: (2025)
AIPO: Improving Training Objective for Iterative Preference Optimization
by: Shen, Yaojie, et al.
Published: (2024)
by: Shen, Yaojie, et al.
Published: (2024)
Chain-of-Thought Reasoning Improves Context-Aware Translation with Large Language Models
by: Ataee, Shabnam, et al.
Published: (2025)
by: Ataee, Shabnam, et al.
Published: (2025)
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
by: Lin, Honglin, et al.
Published: (2025)
by: Lin, Honglin, et al.
Published: (2025)
S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs
by: Du, Yanrui, et al.
Published: (2026)
by: Du, Yanrui, et al.
Published: (2026)
Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework
by: Huang, Kerui, et al.
Published: (2025)
by: Huang, Kerui, et al.
Published: (2025)
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
by: Chua, James, et al.
Published: (2024)
by: Chua, James, et al.
Published: (2024)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
by: Jia, Jinghan, et al.
Published: (2026)
by: Jia, Jinghan, et al.
Published: (2026)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
by: Chen, Yangyi, et al.
Published: (2023)
by: Chen, Yangyi, et al.
Published: (2023)
On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
by: Li, Zhaoyi, et al.
Published: (2026)
by: Li, Zhaoyi, et al.
Published: (2026)
Latent Chain-of-Thought for Visual Reasoning
by: Sun, Guohao, et al.
Published: (2025)
by: Sun, Guohao, et al.
Published: (2025)
Rethinking Chain-of-Thought Reasoning for Videos
by: Zhong, Yiwu, et al.
Published: (2025)
by: Zhong, Yiwu, et al.
Published: (2025)
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation
by: Qiu, Yuli, et al.
Published: (2024)
by: Qiu, Yuli, et al.
Published: (2024)
Fractured Chain-of-Thought Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
by: Ye, Jiacheng, et al.
Published: (2024)
by: Ye, Jiacheng, et al.
Published: (2024)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Faithful Logical Reasoning via Symbolic Chain-of-Thought
by: Xu, Jundong, et al.
Published: (2024)
by: Xu, Jundong, et al.
Published: (2024)
Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
by: Xu, Haolei, et al.
Published: (2025)
by: Xu, Haolei, et al.
Published: (2025)
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
by: Tan, Wentao, et al.
Published: (2025)
by: Tan, Wentao, et al.
Published: (2025)
Similar Items
-
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024) -
Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL
by: Liu, Hanbing, et al.
Published: (2025) -
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
by: Zhang, Hongbo, et al.
Published: (2025) -
Self-Harmonized Chain of Thought
by: Jin, Ziqi, et al.
Published: (2024) -
TinyLlama: An Open-Source Small Language Model
by: Zhang, Peiyuan, et al.
Published: (2024)