DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Peng, Yichen, Song, Jyun-Ting, Jung, Siyeol, Liu, Ruofan, Liu, Haiyang, Chu, Xuangeng, Liu, Ruicong, Wu, Erwin, Koike, Hideki, Kitani, Kris
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911551965888512
author Peng, Yichen
Song, Jyun-Ting
Jung, Siyeol
Liu, Ruofan
Liu, Haiyang
Chu, Xuangeng
Liu, Ruicong
Wu, Erwin
Koike, Hideki
Kitani, Kris
author_facet Peng, Yichen
Song, Jyun-Ting
Jung, Siyeol
Liu, Ruofan
Liu, Haiyang
Chu, Xuangeng
Liu, Ruicong
Wu, Erwin
Koike, Hideki
Kitani, Kris
contents Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without considering social context or modeling the mutual dynamics between two people engaging in conversation. We present DyaDiT, a multi-modal diffusion transformer that generates contextually appropriate human motion from dyadic audio signals. Trained on Seamless Interaction Dataset, DyaDiT takes dyadic audio with optional social-context tokens to produce context-appropriate motion. It fuses information from both speakers to capture interaction dynamics, uses a motion dictionary to encode motion priors, and can optionally utilize the conversational partner's gestures to produce more responsive motion. We evaluate DyaDiT on standard motion generation metrics and conduct quantitative user studies, demonstrating that it not only surpasses existing methods on objective metrics but is also strongly preferred by users, highlighting its robustness and socially favorable motion generation. Code and models will be released upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2602_23165
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
Peng, Yichen
Song, Jyun-Ting
Jung, Siyeol
Liu, Ruofan
Liu, Haiyang
Chu, Xuangeng
Liu, Ruicong
Wu, Erwin
Koike, Hideki
Kitani, Kris
Computer Vision and Pattern Recognition
Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without considering social context or modeling the mutual dynamics between two people engaging in conversation. We present DyaDiT, a multi-modal diffusion transformer that generates contextually appropriate human motion from dyadic audio signals. Trained on Seamless Interaction Dataset, DyaDiT takes dyadic audio with optional social-context tokens to produce context-appropriate motion. It fuses information from both speakers to capture interaction dynamics, uses a motion dictionary to encode motion priors, and can optionally utilize the conversational partner's gestures to produce more responsive motion. We evaluate DyaDiT on standard motion generation metrics and conduct quantitative user studies, demonstrating that it not only surpasses existing methods on objective metrics but is also strongly preferred by users, highlighting its robustness and socially favorable motion generation. Code and models will be released upon acceptance.
title DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.23165