MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pappa, Massimiliano, Collorone, Luca, Ficarra, Giovanni, Spinelli, Indro, Galasso, Fabio
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910436649074688
author Pappa, Massimiliano
Collorone, Luca
Ficarra, Giovanni
Spinelli, Indro
Galasso, Fabio
author_facet Pappa, Massimiliano
Collorone, Luca
Ficarra, Giovanni
Spinelli, Indro
Galasso, Fabio
contents Diffusion Models have revolutionized the field of human motion generation by offering exceptional generation quality and fine-grained controllability through natural language conditioning. Their inherent stochasticity, that is the ability to generate various outputs from a single input, is key to their success. However, this diversity should not be unrestricted, as it may lead to unlikely generations. Instead, it should be confined within the boundaries of text-aligned and realistic generations. To address this issue, we propose MoDiPO (Motion Diffusion DPO), a novel methodology that leverages Direct Preference Optimization (DPO) to align text-to-motion models. We streamline the laborious and expensive process of gathering human preferences needed in DPO by leveraging AI feedback instead. This enables us to experiment with novel DPO strategies, using both online and offline generated motion-preference pairs. To foster future research we contribute with a motion-preference dataset which we dub Pick-a-Move. We demonstrate, both qualitatively and quantitatively, that our proposed method yields significantly more realistic motions. In particular, MoDiPO substantially improves Frechet Inception Distance (FID) while retaining the same RPrecision and Multi-Modality performances.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03803
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
Pappa, Massimiliano
Collorone, Luca
Ficarra, Giovanni
Spinelli, Indro
Galasso, Fabio
Computer Vision and Pattern Recognition
Diffusion Models have revolutionized the field of human motion generation by offering exceptional generation quality and fine-grained controllability through natural language conditioning. Their inherent stochasticity, that is the ability to generate various outputs from a single input, is key to their success. However, this diversity should not be unrestricted, as it may lead to unlikely generations. Instead, it should be confined within the boundaries of text-aligned and realistic generations. To address this issue, we propose MoDiPO (Motion Diffusion DPO), a novel methodology that leverages Direct Preference Optimization (DPO) to align text-to-motion models. We streamline the laborious and expensive process of gathering human preferences needed in DPO by leveraging AI feedback instead. This enables us to experiment with novel DPO strategies, using both online and offline generated motion-preference pairs. To foster future research we contribute with a motion-preference dataset which we dub Pick-a-Move. We demonstrate, both qualitatively and quantitatively, that our proposed method yields significantly more realistic motions. In particular, MoDiPO substantially improves Frechet Inception Distance (FID) while retaining the same RPrecision and Multi-Modality performances.
title MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.03803