Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandit
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Tian, Wang, Shengbo, Li, Ke |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
Online Clustering of Dueling Bandits
by: Wang, Zhiyong, et al.
Published: (2025)
by: Wang, Zhiyong, et al.
Published: (2025)
Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026)
by: Wang, Xiangyi, et al.
Published: (2026)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Dynamic Detection of Relevant Objectives and Adaptation to Preference Drifts in Interactive Evolutionary Multi-Objective Optimization
by: Shavarani, Seyed Mahdi, et al.
Published: (2024)
by: Shavarani, Seyed Mahdi, et al.
Published: (2024)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
Preference-Driven Multi-Objective Combinatorial Optimization with Conditional Computation
by: Fan, Mingfeng, et al.
Published: (2025)
by: Fan, Mingfeng, et al.
Published: (2025)
Analyzing and Overcoming Local Optima in Complex Multi-Objective Optimization by Decomposition-Based Evolutionary Algorithms
by: Dong, Ting, et al.
Published: (2024)
by: Dong, Ting, et al.
Published: (2024)
Forward versus Backward: Comparing Reasoning Objectives in Direct Preference Optimization
by: Nikzad, Murtaza, et al.
Published: (2026)
by: Nikzad, Murtaza, et al.
Published: (2026)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
by: Duan, Shaohua, et al.
Published: (2025)
by: Duan, Shaohua, et al.
Published: (2025)
Expensive Multi-Objective Bayesian Optimization Based on Diffusion Models
by: Li, Bingdong, et al.
Published: (2024)
by: Li, Bingdong, et al.
Published: (2024)
Preference-Guided Diffusion for Multi-Objective Offline Optimization
by: Annadani, Yashas, et al.
Published: (2025)
by: Annadani, Yashas, et al.
Published: (2025)
Offline Multi-Objective Optimization
by: Xue, Ke, et al.
Published: (2024)
by: Xue, Ke, et al.
Published: (2024)
Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching
by: Jin, Jiahuan, et al.
Published: (2026)
by: Jin, Jiahuan, et al.
Published: (2026)
SDPO: Segment-Level Direct Preference Optimization for Social Agents
by: Kong, Aobo, et al.
Published: (2025)
by: Kong, Aobo, et al.
Published: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Token-Importance Guided Direct Preference Optimization
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
Meta-Learning Objectives for Preference Optimization
by: Alfano, Carlo, et al.
Published: (2024)
by: Alfano, Carlo, et al.
Published: (2024)
Interactive Hyperparameter Optimization in Multi-Objective Problems via Preference Learning
by: Giovanelli, Joseph, et al.
Published: (2023)
by: Giovanelli, Joseph, et al.
Published: (2023)
Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
Autoregressive Direct Preference Optimization
by: Oi, Masanari, et al.
Published: (2026)
by: Oi, Masanari, et al.
Published: (2026)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
by: Guo, Yiju, et al.
Published: (2024)
by: Guo, Yiju, et al.
Published: (2024)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
by: Chidambaram, Keertana, et al.
Published: (2025)
by: Chidambaram, Keertana, et al.
Published: (2025)
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
by: Xu, Wenzhe, et al.
Published: (2026)
by: Xu, Wenzhe, et al.
Published: (2026)
In-Context Multi-Objective Optimization
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
EvolMathEval: Towards Evolvable Benchmarks for Mathematical Reasoning via Evolutionary Testing
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
Multi-Armed Bandits-Based Optimization of Decision Trees
by: Shanto, Hasibul Karim, et al.
Published: (2025)
by: Shanto, Hasibul Karim, et al.
Published: (2025)
Preference-Conditioned Gradient Variations for Multi-Objective Quality-Diversity
by: Janmohamed, Hannah, et al.
Published: (2024)
by: Janmohamed, Hannah, et al.
Published: (2024)
Unlearning of Knowledge Graph Embedding via Preference Optimization
by: Liu, Jiajun, et al.
Published: (2025)
by: Liu, Jiajun, et al.
Published: (2025)
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
by: Ambadkar, Tanmay, et al.
Published: (2026)
by: Ambadkar, Tanmay, et al.
Published: (2026)
Token-level Direct Preference Optimization
by: Zeng, Yongcheng, et al.
Published: (2024)
by: Zeng, Yongcheng, et al.
Published: (2024)
An Inverse Modeling Constrained Multi-Objective Evolutionary Algorithm Based on Decomposition
by: Farias, Lucas R. C., et al.
Published: (2024)
by: Farias, Lucas R. C., et al.
Published: (2024)
Rank-Based Learning and Local Model Based Evolutionary Algorithm for High-Dimensional Expensive Multi-Objective Problems
by: Chen, Guodong, et al.
Published: (2023)
by: Chen, Guodong, et al.
Published: (2023)
MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning
by: Lin, Yunze
Published: (2025)
by: Lin, Yunze
Published: (2025)
On Softmax Direct Preference Optimization for Recommendation
by: Chen, Yuxin, et al.
Published: (2024)
by: Chen, Yuxin, et al.
Published: (2024)
Archive-based Single-Objective Evolutionary Algorithms for Submodular Optimization
by: Neumann, Frank, et al.
Published: (2024)
by: Neumann, Frank, et al.
Published: (2024)
Similar Items
-
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024) -
Online Clustering of Dueling Bandits
by: Wang, Zhiyong, et al.
Published: (2025) -
Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback
by: Wang, Shengbo, et al.
Published: (2025) -
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026) -
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)