Saved in:
| Main Authors: | Zhang, Lijun, Li, Lin, Qi, Yajie, Song, Huizhong, Yang, Yaodong, Wang, Jun, Wei, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.20359 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Constrained Language Model Policy Optimization via Risk-aware Stepwise Alignment
by: Zhang, Lijun, et al.
Published: (2025)
by: Zhang, Lijun, et al.
Published: (2025)
Goal-Guided Efficient Exploration via Large Language Model in Reinforcement Learning
by: Qi, Yajie, et al.
Published: (2025)
by: Qi, Yajie, et al.
Published: (2025)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Towards Unsupervised Training of Matching-based Graph Edit Distance Solver via Preference-aware GAN
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Orthogonal Finetuning for Direct Preference Optimization
by: Yang, Chenxu, et al.
Published: (2024)
by: Yang, Chenxu, et al.
Published: (2024)
MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning
by: Lin, Yunze
Published: (2025)
by: Lin, Yunze
Published: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Graph Unlearning Meets Influence-aware Negative Preference Optimization
by: Chen, Qiang, et al.
Published: (2025)
by: Chen, Qiang, et al.
Published: (2025)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
by: Jiang, Zaifan, et al.
Published: (2023)
by: Jiang, Zaifan, et al.
Published: (2023)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
Continuous-Utility Direct Preference Optimization
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2026)
Diffusion-Modeled Reinforcement Learning for Carbon and Risk-Aware Microgrid Optimization
by: Zhao, Yunyi, et al.
Published: (2025)
by: Zhao, Yunyi, et al.
Published: (2025)
Contextual Preference Collaborative Measure Framework Based on Belief System
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk
by: Woydt, Tim, et al.
Published: (2026)
by: Woydt, Tim, et al.
Published: (2026)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
by: Ren, Tao, et al.
Published: (2025)
by: Ren, Tao, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
A Unified Gaussian Process for Branching and Nested Hyperparameter Optimization
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
by: Wang, Junzhe, et al.
Published: (2026)
by: Wang, Junzhe, et al.
Published: (2026)
Emergent Risk Awareness in Rational Agents under Resource Constraints
by: Ornia, Daniel Jarne, et al.
Published: (2025)
by: Ornia, Daniel Jarne, et al.
Published: (2025)
Hindsight Preference Optimization for Financial Time Series Advisory
by: Cui, Yanwei, et al.
Published: (2026)
by: Cui, Yanwei, et al.
Published: (2026)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
Thinking Preference Optimization
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Selective Conformal Risk Control
by: Xu, Yunpeng, et al.
Published: (2025)
by: Xu, Yunpeng, et al.
Published: (2025)
Uncertainty-aware Traffic Prediction under Missing Data
by: Mei, Hao, et al.
Published: (2023)
by: Mei, Hao, et al.
Published: (2023)
Generative Risk Minimization for Out-of-Distribution Generalization on Graphs
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
by: Shen, Xiangwei, et al.
Published: (2025)
by: Shen, Xiangwei, et al.
Published: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Causal-aware Graph Neural Architecture Search under Distribution Shifts
by: Li, Peiwen, et al.
Published: (2024)
by: Li, Peiwen, et al.
Published: (2024)
MO-RiskVAE: A Multi-Omics Variational Autoencoder for Survival Risk Modeling in Multiple MyelomaMO-RiskVAE
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
Similar Items
-
Constrained Language Model Policy Optimization via Risk-aware Stepwise Alignment
by: Zhang, Lijun, et al.
Published: (2025) -
Goal-Guided Efficient Exploration via Large Language Model in Reinforcement Learning
by: Qi, Yajie, et al.
Published: (2025) -
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
by: Wei, Zeming, et al.
Published: (2026) -
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024) -
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)