ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909433919963136 |
|---|---|
| author | Lu, Shunlin Wang, Jingbo Lu, Zeyu Chen, Ling-Hao Dai, Wenxun Dong, Junting Dou, Zhiyang Dai, Bo Zhang, Ruimao |
| author_facet | Lu, Shunlin Wang, Jingbo Lu, Zeyu Chen, Ling-Hao Dai, Wenxun Dong, Junting Dou, Zhiyang Dai, Bo Zhang, Ruimao |
| contents | The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion tokenizer Motion FSQ-VAE and a text-prefix autoregressive transformer. Through comprehensive experiments, we observe the scaling behavior of this system. For the first time, we confirm the existence of scaling laws within the context of motion generation. Specifically, our results demonstrate that the normalized test loss of our prefix autoregressive models adheres to a logarithmic law in relation to compute budgets. Furthermore, we also confirm the power law between Non-Vocabulary Parameters, Vocabulary Parameters, and Data Tokens with respect to compute budgets respectively. Leveraging the scaling law, we predict the optimal transformer size, vocabulary size, and data requirements for a compute budget of $1e18$. The test loss of the system, when trained with the optimal model size, vocabulary size, and required data, aligns precisely with the predicted test loss, thereby validating the scaling law. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_14559 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model Lu, Shunlin Wang, Jingbo Lu, Zeyu Chen, Ling-Hao Dai, Wenxun Dong, Junting Dou, Zhiyang Dai, Bo Zhang, Ruimao Computer Vision and Pattern Recognition Machine Learning The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion tokenizer Motion FSQ-VAE and a text-prefix autoregressive transformer. Through comprehensive experiments, we observe the scaling behavior of this system. For the first time, we confirm the existence of scaling laws within the context of motion generation. Specifically, our results demonstrate that the normalized test loss of our prefix autoregressive models adheres to a logarithmic law in relation to compute budgets. Furthermore, we also confirm the power law between Non-Vocabulary Parameters, Vocabulary Parameters, and Data Tokens with respect to compute budgets respectively. Leveraging the scaling law, we predict the optimal transformer size, vocabulary size, and data requirements for a compute budget of $1e18$. The test loss of the system, when trained with the optimal model size, vocabulary size, and required data, aligns precisely with the predicted test loss, thereby validating the scaling law. |
| title | ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2412.14559 |