Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yuheng, Yu, Dian, Ge, Tao, Song, Linfeng, Zeng, Zhichen, Mi, Haitao, Jiang, Nan, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
von: Yu, Dian, et al.
Veröffentlicht: (2024)
von: Yu, Dian, et al.
Veröffentlicht: (2024)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
Teaching LLMs to Refine with Tools
von: Yu, Dian, et al.
Veröffentlicht: (2024)
von: Yu, Dian, et al.
Veröffentlicht: (2024)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
von: Yu, Dian, et al.
Veröffentlicht: (2025)
von: Yu, Dian, et al.
Veröffentlicht: (2025)
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
LiteSearch: Efficacious Tree Search for LLM
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
von: Wang, Ante, et al.
Veröffentlicht: (2025)
von: Wang, Ante, et al.
Veröffentlicht: (2025)
Scaling Synthetic Data Creation with 1,000,000,000 Personas
von: Ge, Tao, et al.
Veröffentlicht: (2024)
von: Ge, Tao, et al.
Veröffentlicht: (2024)
Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
von: Das, Souvik, et al.
Veröffentlicht: (2024)
von: Das, Souvik, et al.
Veröffentlicht: (2024)
Verified Critical Step Optimization for LLM Agents
von: Li, Mukai, et al.
Veröffentlicht: (2026)
von: Li, Mukai, et al.
Veröffentlicht: (2026)
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation
von: Li, Xiangci, et al.
Veröffentlicht: (2024)
von: Li, Xiangci, et al.
Veröffentlicht: (2024)
Inconsistent dialogue responses and how to recover from them
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
One Token to Fool LLM-as-a-Judge
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
Optimistic Online Mirror Descent for Bridging Stochastic and Adversarial Online Convex Optimization
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
Collaborative decoding of critical tokens for boosting factuality of large language models
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows
von: Yao, Wenlin, et al.
Veröffentlicht: (2024)
von: Yao, Wenlin, et al.
Veröffentlicht: (2024)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
von: Xu, Hang, et al.
Veröffentlicht: (2024)
von: Xu, Hang, et al.
Veröffentlicht: (2024)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
Preference Alignment Improves Language Model-Based TTS
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
HunyuanProver: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2025)
EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
von: Li, Mukai, et al.
Veröffentlicht: (2025)
von: Li, Mukai, et al.
Veröffentlicht: (2025)
Self-Consistency Boosts Calibration for Math Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Multilingual Blending: LLM Safety Alignment Evaluation with Language Mixture
von: Song, Jiayang, et al.
Veröffentlicht: (2024)
von: Song, Jiayang, et al.
Veröffentlicht: (2024)
Less is More: Improving LLM Alignment via Preference Data Selection
von: Deng, Xun, et al.
Veröffentlicht: (2025)
von: Deng, Xun, et al.
Veröffentlicht: (2025)
PABU: Progress-Aware Belief Update for Efficient LLM Agents
von: Jiang, Haitao, et al.
Veröffentlicht: (2026)
von: Jiang, Haitao, et al.
Veröffentlicht: (2026)
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024) -
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
von: Yu, Dian, et al.
Veröffentlicht: (2024) -
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
von: Wang, Xiyao, et al.
Veröffentlicht: (2024) -
Teaching LLMs to Refine with Tools
von: Yu, Dian, et al.
Veröffentlicht: (2024) -
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
von: Yu, Dian, et al.
Veröffentlicht: (2025)