InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Dongwei, Guo, Longwei, Guan, Jiazhi, Huang, Luying, Li, Yiding, Liu, Haojie, Feng, Haocheng, He, Wei, Wang, Kaisiyuan, Zhou, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
by: Guan, Jiazhi, et al.
Published: (2026)
by: Guan, Jiazhi, et al.
Published: (2026)
ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration
by: Yang, Fengyuan, et al.
Published: (2026)
by: Yang, Fengyuan, et al.
Published: (2026)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
by: Yang, Quanwei, et al.
Published: (2025)
by: Yang, Quanwei, et al.
Published: (2025)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
by: Shen, Zhelun, et al.
Published: (2025)
by: Shen, Zhelun, et al.
Published: (2025)
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
by: Fan, Yingying, et al.
Published: (2025)
by: Fan, Yingying, et al.
Published: (2025)
TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
by: Tong, Jinguang, et al.
Published: (2026)
by: Tong, Jinguang, et al.
Published: (2026)
AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation
by: Sun, Yasheng, et al.
Published: (2024)
by: Sun, Yasheng, et al.
Published: (2024)
GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
by: Huang, Xuan, et al.
Published: (2026)
by: Huang, Xuan, et al.
Published: (2026)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
by: Xu, Anfeng, et al.
Published: (2024)
by: Xu, Anfeng, et al.
Published: (2024)
Proactive Recommendation with Iterative Preference Guidance
by: Bi, Shuxian, et al.
Published: (2024)
by: Bi, Shuxian, et al.
Published: (2024)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
by: Herbuela, Von Ralph Dane Marquez, et al.
Published: (2025)
by: Herbuela, Von Ralph Dane Marquez, et al.
Published: (2025)
Inter-Stance: A Dyadic Multimodal Corpus for Conversational Stance Analysis
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
InterChat: Enhancing Generative Visual Analytics using Multimodal Interactions
by: Chen, Juntong, et al.
Published: (2025)
by: Chen, Juntong, et al.
Published: (2025)
InterChat: Enhancing Generative Visual Analytics using Multimodal Interactions
by: Juntong Chen, et al.
Published: (2025)
by: Juntong Chen, et al.
Published: (2025)
Spatiotemporal Emotional Synchrony in Dyadic Interactions: The Role of Speech Conditions in Facial and Vocal Affective Alignment
by: Herbuela, Von Ralph Dane Marquez, et al.
Published: (2025)
by: Herbuela, Von Ralph Dane Marquez, et al.
Published: (2025)
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
by: Akkerman, Rick, et al.
Published: (2024)
by: Akkerman, Rick, et al.
Published: (2024)
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
by: Xu, Longwei, et al.
Published: (2026)
by: Xu, Longwei, et al.
Published: (2026)
DiffusionCounterfactuals: Inferring High-dimensional Counterfactuals with Guidance of Causal Representations
by: Zhu, Jiageng, et al.
Published: (2024)
by: Zhu, Jiageng, et al.
Published: (2024)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
Visual Generation Without Guidance
by: Chen, Huayu, et al.
Published: (2025)
by: Chen, Huayu, et al.
Published: (2025)
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
by: Liao, Xinyao, et al.
Published: (2026)
by: Liao, Xinyao, et al.
Published: (2026)
Dyadic Interaction Modeling for Social Behavior Generation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Effect of a DBT‐Informed Dyadic Intervention (Better Together) on Psychosocial Outcomes in Colorectal Cancer Patient–Spouse Dyads: A Randomized Controlled Trial
by: Yanfei Jin, et al.
Published: (2026)
by: Yanfei Jin, et al.
Published: (2026)
Intermediate dimensions of Moran sets and their visualization
by: Du, Yali, et al.
Published: (2024)
by: Du, Yali, et al.
Published: (2024)
Chapter Exploring the Dynamics of Teachers' Technology Adoption through a Longitudinal Lens
by: Zheng, Longwei
Published: (2026)
by: Zheng, Longwei
Published: (2026)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
by: Wang, He, et al.
Published: (2024)
by: Wang, He, et al.
Published: (2024)
StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
by: Jeong, Jaeseok, et al.
Published: (2025)
by: Jeong, Jaeseok, et al.
Published: (2025)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
by: Jia, Zhenqi, et al.
Published: (2024)
by: Jia, Zhenqi, et al.
Published: (2024)
Combinatorial Analysis of Dyadic and Quasi-Dyadic Codes
by: Gómez-Fonseca, Anthony, et al.
Published: (2026)
by: Gómez-Fonseca, Anthony, et al.
Published: (2026)
GlobalPaint: Spatiotemporal Coherent Video Outpainting with Global Feature Guidance
by: Pan, Yueming, et al.
Published: (2026)
by: Pan, Yueming, et al.
Published: (2026)
Gaussian process learning of nonlinear dynamics
by: Ye, Dongwei, et al.
Published: (2023)
by: Ye, Dongwei, et al.
Published: (2023)
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023
by: Wang, He, et al.
Published: (2024)
by: Wang, He, et al.
Published: (2024)
Inter‐prefectural regional disparities in gastric cancer surgery: A Japanese nationwide population‐based cohort study from 2014 to 2019
by: Masamitsu Kido, et al.
Published: (2024)
by: Masamitsu Kido, et al.
Published: (2024)
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
by: Lu, Haocheng, et al.
Published: (2026)
by: Lu, Haocheng, et al.
Published: (2026)
Similar Items
-
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
by: Guan, Jiazhi, et al.
Published: (2026) -
ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration
by: Yang, Fengyuan, et al.
Published: (2026) -
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025) -
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
by: Yang, Quanwei, et al.
Published: (2025) -
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
by: Guan, Jiazhi, et al.
Published: (2024)