DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guan, Jiazhi, Yang, Quanwei, Huang, Luying, Liang, Junhao, Liang, Borong, Feng, Haocheng, He, Wei, Wang, Kaisiyuan, Zhou, Hang, Wang, Jingdong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915851251220480
author Guan, Jiazhi
Yang, Quanwei
Huang, Luying
Liang, Junhao
Liang, Borong
Feng, Haocheng
He, Wei
Wang, Kaisiyuan
Zhou, Hang
Wang, Jingdong
author_facet Guan, Jiazhi
Yang, Quanwei
Huang, Luying
Liang, Junhao
Liang, Borong
Feng, Haocheng
He, Wei
Wang, Kaisiyuan
Zhou, Hang
Wang, Jingdong
contents Human-centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human-Object Interaction (HOI) videos. Existing works rely on dense control signals, template videos, or carefully crafted text prompts, which limit flexibility and generalization to novel objects. We introduce a framework, namely DISPLAY, guided by Sparse Motion Guidance, composed only of wrist joint coordinates and a shape-agnostic object bounding box. This lightweight guidance alleviates the imbalance between human and object representations and enables intuitive user control. To enhance fidelity under such sparse conditions, we propose an Object-Stressed Attention mechanism that improves object robustness. To address the scarcity of high-quality HOI data, we further develop a Multi-Task Auxiliary Training strategy with a dedicated data curation pipeline, allowing the model to benefit from both reliable HOI samples and auxiliary tasks. Comprehensive experiments show that our method achieves high-fidelity, controllable HOI generation across diverse tasks. The project page can be found at \href{https://mumuwei.github.io/DISPLAY/}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09883
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
Guan, Jiazhi
Yang, Quanwei
Huang, Luying
Liang, Junhao
Liang, Borong
Feng, Haocheng
He, Wei
Wang, Kaisiyuan
Zhou, Hang
Wang, Jingdong
Computer Vision and Pattern Recognition
Human-centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human-Object Interaction (HOI) videos. Existing works rely on dense control signals, template videos, or carefully crafted text prompts, which limit flexibility and generalization to novel objects. We introduce a framework, namely DISPLAY, guided by Sparse Motion Guidance, composed only of wrist joint coordinates and a shape-agnostic object bounding box. This lightweight guidance alleviates the imbalance between human and object representations and enables intuitive user control. To enhance fidelity under such sparse conditions, we propose an Object-Stressed Attention mechanism that improves object robustness. To address the scarcity of high-quality HOI data, we further develop a Multi-Task Auxiliary Training strategy with a dedicated data curation pipeline, allowing the model to benefit from both reliable HOI samples and auxiliary tasks. Comprehensive experiments show that our method achieves high-fidelity, controllable HOI generation across diverse tasks. The project page can be found at \href{https://mumuwei.github.io/DISPLAY/}.
title DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.09883