ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lu, Shunlin, Wang, Jingbo, Lu, Zeyu, Chen, Ling-Hao, Dai, Wenxun, Dong, Junting, Dou, Zhiyang, Dai, Bo, Zhang, Ruimao
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909433919963136
author Lu, Shunlin
Wang, Jingbo
Lu, Zeyu
Chen, Ling-Hao
Dai, Wenxun
Dong, Junting
Dou, Zhiyang
Dai, Bo
Zhang, Ruimao
author_facet Lu, Shunlin
Wang, Jingbo
Lu, Zeyu
Chen, Ling-Hao
Dai, Wenxun
Dong, Junting
Dou, Zhiyang
Dai, Bo
Zhang, Ruimao
contents The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion tokenizer Motion FSQ-VAE and a text-prefix autoregressive transformer. Through comprehensive experiments, we observe the scaling behavior of this system. For the first time, we confirm the existence of scaling laws within the context of motion generation. Specifically, our results demonstrate that the normalized test loss of our prefix autoregressive models adheres to a logarithmic law in relation to compute budgets. Furthermore, we also confirm the power law between Non-Vocabulary Parameters, Vocabulary Parameters, and Data Tokens with respect to compute budgets respectively. Leveraging the scaling law, we predict the optimal transformer size, vocabulary size, and data requirements for a compute budget of $1e18$. The test loss of the system, when trained with the optimal model size, vocabulary size, and required data, aligns precisely with the predicted test loss, thereby validating the scaling law.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14559
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
Lu, Shunlin
Wang, Jingbo
Lu, Zeyu
Chen, Ling-Hao
Dai, Wenxun
Dong, Junting
Dou, Zhiyang
Dai, Bo
Zhang, Ruimao
Computer Vision and Pattern Recognition
Machine Learning
The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion tokenizer Motion FSQ-VAE and a text-prefix autoregressive transformer. Through comprehensive experiments, we observe the scaling behavior of this system. For the first time, we confirm the existence of scaling laws within the context of motion generation. Specifically, our results demonstrate that the normalized test loss of our prefix autoregressive models adheres to a logarithmic law in relation to compute budgets. Furthermore, we also confirm the power law between Non-Vocabulary Parameters, Vocabulary Parameters, and Data Tokens with respect to compute budgets respectively. Leveraging the scaling law, we predict the optimal transformer size, vocabulary size, and data requirements for a compute budget of $1e18$. The test loss of the system, when trained with the optimal model size, vocabulary size, and required data, aligns precisely with the predicted test loss, thereby validating the scaling law.
title ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.14559