Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914556734865408 |
|---|---|
| author | Longhini, Alberta Emukpere, David Renders, Jean-Michel Kim, Seungsu |
| author_facet | Longhini, Alberta Emukpere, David Renders, Jean-Michel Kim, Seungsu |
| contents | We address the problem of fine-tuning pre-trained generative policies with reinforcement learning (RL) while preserving the multimodality of their action distributions. Existing methods for RL fine-tuning of generative policies (e.g., diffusion policies) improve task performance but often collapse diverse behaviors into a single reward-maximizing mode. To mitigate this issue, we propose an unsupervised mode discovery framework that uncovers latent behavioral modes within generative policies. The discovered modes enable the use of mutual information as an intrinsic reward, regularizing RL fine-tuning to enhance task success while maintaining behavioral diversity. Experiments on robotic manipulation tasks demonstrate that our method consistently outperforms conventional fine-tuning approaches, achieving higher success rates and preserving richer multimodal action distributions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11387 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies Longhini, Alberta Emukpere, David Renders, Jean-Michel Kim, Seungsu Machine Learning Robotics We address the problem of fine-tuning pre-trained generative policies with reinforcement learning (RL) while preserving the multimodality of their action distributions. Existing methods for RL fine-tuning of generative policies (e.g., diffusion policies) improve task performance but often collapse diverse behaviors into a single reward-maximizing mode. To mitigate this issue, we propose an unsupervised mode discovery framework that uncovers latent behavioral modes within generative policies. The discovered modes enable the use of mutual information as an intrinsic reward, regularizing RL fine-tuning to enhance task success while maintaining behavioral diversity. Experiments on robotic manipulation tasks demonstrate that our method consistently outperforms conventional fine-tuning approaches, achieving higher success rates and preserving richer multimodal action distributions. |
| title | Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies |
| topic | Machine Learning Robotics |
| url | https://arxiv.org/abs/2605.11387 |