SocialGen: Modeling Multi-Human Social Interaction with Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909557399224320 |
|---|---|
| author | Yu, Heng Zhang, Juze Chen, Changan Xiang, Tiange Fang, Yusu Niebles, Juan Carlos Adeli, Ehsan |
| author_facet | Yu, Heng Zhang, Juze Chen, Changan Xiang, Tiange Fang, Yusu Niebles, Juan Carlos Adeli, Ehsan |
| contents | Human interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental to a wide range of real-world applications. In this paper, we introduce SocialGen, the first unified motion-language model capable of modeling interaction behaviors among varying numbers of individuals, to address this crucial yet challenging problem. Unlike prior methods that are limited to two-person interactions, we propose a novel social motion representation that supports tokenizing the motions of an arbitrary number of individuals and aligning them with the language space. This alignment enables the model to leverage rich, pretrained linguistic knowledge to better understand and reason about human social behaviors. To tackle the challenges of data scarcity, we curate a comprehensive multi-human interaction dataset, SocialX, enriched with textual annotations. Leveraging this dataset, we establish the first comprehensive benchmark for multi-human interaction tasks. Our method achieves state-of-the-art performance across motion-language tasks, setting a new standard for multi-human interaction modeling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_22906 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SocialGen: Modeling Multi-Human Social Interaction with Language Models Yu, Heng Zhang, Juze Chen, Changan Xiang, Tiange Fang, Yusu Niebles, Juan Carlos Adeli, Ehsan Computer Vision and Pattern Recognition Human interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental to a wide range of real-world applications. In this paper, we introduce SocialGen, the first unified motion-language model capable of modeling interaction behaviors among varying numbers of individuals, to address this crucial yet challenging problem. Unlike prior methods that are limited to two-person interactions, we propose a novel social motion representation that supports tokenizing the motions of an arbitrary number of individuals and aligning them with the language space. This alignment enables the model to leverage rich, pretrained linguistic knowledge to better understand and reason about human social behaviors. To tackle the challenges of data scarcity, we curate a comprehensive multi-human interaction dataset, SocialX, enriched with textual annotations. Leveraging this dataset, we establish the first comprehensive benchmark for multi-human interaction tasks. Our method achieves state-of-the-art performance across motion-language tasks, setting a new standard for multi-human interaction modeling. |
| title | SocialGen: Modeling Multi-Human Social Interaction with Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.22906 |