Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Tian, Wang, Shengbo, Li, Ke
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916645975359488
author Huang, Tian
Wang, Shengbo
Li, Ke
author_facet Huang, Tian
Wang, Shengbo
Li, Ke
contents Optimization problems find widespread use in both single-objective and multi-objective scenarios. In practical applications, users aspire for solutions that converge to the region of interest (ROI) along the Pareto front (PF). While the conventional approach involves approximating a fitness function or an objective function to reflect user preferences, this paper explores an alternative avenue. Specifically, we aim to discover a method that sidesteps the need for calculating the fitness function, relying solely on human feedback. Our proposed approach entails conducting direct preference learning facilitated by an active dueling bandit algorithm. The experimental phase is structured into three sessions. Firstly, we assess the performance of our active dueling bandit algorithm. Secondly, we implement our proposed method within the context of Multi-objective Evolutionary Algorithms (MOEAs). Finally, we deploy our method in a practical problem, specifically in protein structure prediction (PSP). This research presents a novel interactive preference-based MOEA framework that not only addresses the limitations of traditional techniques but also unveils new possibilities for optimization problems.
format Preprint
id arxiv_https___arxiv_org_abs_2311_14003
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandit
Huang, Tian
Wang, Shengbo
Li, Ke
Artificial Intelligence
Optimization problems find widespread use in both single-objective and multi-objective scenarios. In practical applications, users aspire for solutions that converge to the region of interest (ROI) along the Pareto front (PF). While the conventional approach involves approximating a fitness function or an objective function to reflect user preferences, this paper explores an alternative avenue. Specifically, we aim to discover a method that sidesteps the need for calculating the fitness function, relying solely on human feedback. Our proposed approach entails conducting direct preference learning facilitated by an active dueling bandit algorithm. The experimental phase is structured into three sessions. Firstly, we assess the performance of our active dueling bandit algorithm. Secondly, we implement our proposed method within the context of Multi-objective Evolutionary Algorithms (MOEAs). Finally, we deploy our method in a practical problem, specifically in protein structure prediction (PSP). This research presents a novel interactive preference-based MOEA framework that not only addresses the limitations of traditional techniques but also unveils new possibilities for optimization problems.
title Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandit
topic Artificial Intelligence
url https://arxiv.org/abs/2311.14003