MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yin, li, Mu, Leng, Zhiying, Li, Frederick W. B., Liang, Xiaohui
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913933405716480
author Wang, Yin
li, Mu
Leng, Zhiying
Li, Frederick W. B.
Liang, Xiaohui
author_facet Wang, Yin
li, Mu
Leng, Zhiying
Li, Frederick W. B.
Liang, Xiaohui
contents We introduce MOST, a novel motion diffusion model via temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction
Wang, Yin
li, Mu
Leng, Zhiying
Li, Frederick W. B.
Liang, Xiaohui
Computer Vision and Pattern Recognition
We introduce MOST, a novel motion diffusion model via temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts.
title MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.06590