FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Woosung, Lee, Jinho, Lee, Jongmin, Lee, Byung-Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917087478284288
author Kim, Woosung
Lee, Jinho
Lee, Jongmin
Lee, Byung-Jun
author_facet Kim, Woosung
Lee, Jinho
Lee, Jongmin
Lee, Byung-Jun
contents Multi-objective reinforcement learning (MORL) aims to optimize policies in the presence of conflicting objectives, where linear scalarization is commonly used to reduce vector-valued returns into scalar signals. While effective for certain preferences, this approach cannot capture fairness-oriented goals such as Nash social welfare or max-min fairness, which require nonlinear and non-additive trade-offs. Although several online algorithms have been proposed for specific fairness objectives, a unified approach for optimizing nonlinear welfare criteria in the offline setting-where learning must proceed from a fixed dataset-remains unexplored. In this work, we present FairDICE, the first offline MORL framework that directly optimizes nonlinear welfare objective. FairDICE leverages distribution correction estimation to jointly account for welfare maximization and distributional regularization, enabling stable and sample-efficient learning without requiring explicit preference weights or exhaustive weight search. Across multiple offline benchmarks, FairDICE demonstrates strong fairness-aware performance compared to existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
Kim, Woosung
Lee, Jinho
Lee, Jongmin
Lee, Byung-Jun
Machine Learning
Artificial Intelligence
Multi-objective reinforcement learning (MORL) aims to optimize policies in the presence of conflicting objectives, where linear scalarization is commonly used to reduce vector-valued returns into scalar signals. While effective for certain preferences, this approach cannot capture fairness-oriented goals such as Nash social welfare or max-min fairness, which require nonlinear and non-additive trade-offs. Although several online algorithms have been proposed for specific fairness objectives, a unified approach for optimizing nonlinear welfare criteria in the offline setting-where learning must proceed from a fixed dataset-remains unexplored. In this work, we present FairDICE, the first offline MORL framework that directly optimizes nonlinear welfare objective. FairDICE leverages distribution correction estimation to jointly account for welfare maximization and distributional regularization, enabling stable and sample-efficient learning without requiring explicit preference weights or exhaustive weight search. Across multiple offline benchmarks, FairDICE demonstrates strong fairness-aware performance compared to existing baselines.
title FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.08062