Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Longhini, Alberta, Emukpere, David, Renders, Jean-Michel, Kim, Seungsu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914556734865408
author Longhini, Alberta
Emukpere, David
Renders, Jean-Michel
Kim, Seungsu
author_facet Longhini, Alberta
Emukpere, David
Renders, Jean-Michel
Kim, Seungsu
contents We address the problem of fine-tuning pre-trained generative policies with reinforcement learning (RL) while preserving the multimodality of their action distributions. Existing methods for RL fine-tuning of generative policies (e.g., diffusion policies) improve task performance but often collapse diverse behaviors into a single reward-maximizing mode. To mitigate this issue, we propose an unsupervised mode discovery framework that uncovers latent behavioral modes within generative policies. The discovered modes enable the use of mutual information as an intrinsic reward, regularizing RL fine-tuning to enhance task success while maintaining behavioral diversity. Experiments on robotic manipulation tasks demonstrate that our method consistently outperforms conventional fine-tuning approaches, achieving higher success rates and preserving richer multimodal action distributions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11387
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies
Longhini, Alberta
Emukpere, David
Renders, Jean-Michel
Kim, Seungsu
Machine Learning
Robotics
We address the problem of fine-tuning pre-trained generative policies with reinforcement learning (RL) while preserving the multimodality of their action distributions. Existing methods for RL fine-tuning of generative policies (e.g., diffusion policies) improve task performance but often collapse diverse behaviors into a single reward-maximizing mode. To mitigate this issue, we propose an unsupervised mode discovery framework that uncovers latent behavioral modes within generative policies. The discovered modes enable the use of mutual information as an intrinsic reward, regularizing RL fine-tuning to enhance task success while maintaining behavioral diversity. Experiments on robotic manipulation tasks demonstrate that our method consistently outperforms conventional fine-tuning approaches, achieving higher success rates and preserving richer multimodal action distributions.
title Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies
topic Machine Learning
Robotics
url https://arxiv.org/abs/2605.11387