Exploring Text-to-Motion Generation with Human Preference

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sheng, Jenny, Lin, Matthieu, Zhao, Andrew, Pruvost, Kevin, Wen, Yu-Hui, Li, Yangguang, Huang, Gao, Liu, Yong-Jin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911839409930240
author Sheng, Jenny
Lin, Matthieu
Zhao, Andrew
Pruvost, Kevin
Wen, Yu-Hui
Li, Yangguang
Huang, Gao
Liu, Yong-Jin
author_facet Sheng, Jenny
Lin, Matthieu
Zhao, Andrew
Pruvost, Kevin
Wen, Yu-Hui
Li, Yangguang
Huang, Gao
Liu, Yong-Jin
contents This paper presents an exploration of preference learning in text-to-motion generation. We find that current improvements in text-to-motion generation still rely on datasets requiring expert labelers with motion capture systems. Instead, learning from human preference data does not require motion capture systems; a labeler with no expertise simply compares two generated motions. This is particularly efficient because evaluating the model's output is easier than gathering the motion that performs a desired task (e.g. backflip). To pioneer the exploration of this paradigm, we annotate 3,528 preference pairs generated by MotionGPT, marking the first effort to investigate various algorithms for learning from preference data. In particular, our exploration highlights important design choices when using preference data. Additionally, our experimental results show that preference learning has the potential to greatly improve current text-to-motion generative models. Our code and dataset are publicly available at https://github.com/THU-LYJ-Lab/InstructMotion}{https://github.com/THU-LYJ-Lab/InstructMotion to further facilitate research in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09445
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Text-to-Motion Generation with Human Preference
Sheng, Jenny
Lin, Matthieu
Zhao, Andrew
Pruvost, Kevin
Wen, Yu-Hui
Li, Yangguang
Huang, Gao
Liu, Yong-Jin
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
This paper presents an exploration of preference learning in text-to-motion generation. We find that current improvements in text-to-motion generation still rely on datasets requiring expert labelers with motion capture systems. Instead, learning from human preference data does not require motion capture systems; a labeler with no expertise simply compares two generated motions. This is particularly efficient because evaluating the model's output is easier than gathering the motion that performs a desired task (e.g. backflip). To pioneer the exploration of this paradigm, we annotate 3,528 preference pairs generated by MotionGPT, marking the first effort to investigate various algorithms for learning from preference data. In particular, our exploration highlights important design choices when using preference data. Additionally, our experimental results show that preference learning has the potential to greatly improve current text-to-motion generative models. Our code and dataset are publicly available at https://github.com/THU-LYJ-Lab/InstructMotion}{https://github.com/THU-LYJ-Lab/InstructMotion to further facilitate research in this area.
title Exploring Text-to-Motion Generation with Human Preference
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.09445