MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Pang, Youxin, Liu, Jiajun, Tan, Lingfeng, Zhang, Yong, Gao, Feng, Deng, Xiang, Kang, Zhuoliang, Wei, Xiaoming, Liu, Yebin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
di: Deng, Xiang, et al.
Pubblicazione: (2026)
di: Deng, Xiang, et al.
Pubblicazione: (2026)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
di: Pang, Youxin, et al.
Pubblicazione: (2025)
di: Pang, Youxin, et al.
Pubblicazione: (2025)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
di: Shao, Ruizhi, et al.
Pubblicazione: (2024)
di: Shao, Ruizhi, et al.
Pubblicazione: (2024)
Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
di: Deng, Xiang, et al.
Pubblicazione: (2024)
di: Deng, Xiang, et al.
Pubblicazione: (2024)
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
di: Kong, Zhe, et al.
Pubblicazione: (2025)
di: Kong, Zhe, et al.
Pubblicazione: (2025)
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
di: Pang, Youxin, et al.
Pubblicazione: (2024)
di: Pang, Youxin, et al.
Pubblicazione: (2024)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
di: Yang, Shaoshu, et al.
Pubblicazione: (2025)
di: Yang, Shaoshu, et al.
Pubblicazione: (2025)
MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
di: Wang, Qian, et al.
Pubblicazione: (2025)
di: Wang, Qian, et al.
Pubblicazione: (2025)
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
di: Chen, Yushuo, et al.
Pubblicazione: (2025)
di: Chen, Yushuo, et al.
Pubblicazione: (2025)
GMTalker: Gaussian Mixture-based Audio-Driven Emotional Talking Video Portraits
di: Xia, Yibo, et al.
Pubblicazione: (2023)
di: Xia, Yibo, et al.
Pubblicazione: (2023)
WildActor: Unconstrained Identity-Preserving Video Generation
di: Guo, Qin, et al.
Pubblicazione: (2026)
di: Guo, Qin, et al.
Pubblicazione: (2026)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
di: Liu, Xiangyu, et al.
Pubblicazione: (2026)
di: Liu, Xiangyu, et al.
Pubblicazione: (2026)
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling
di: Zhang, Min, et al.
Pubblicazione: (2024)
di: Zhang, Min, et al.
Pubblicazione: (2024)
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2025)
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2025)
RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios
di: Zheng, Liming, et al.
Pubblicazione: (2024)
di: Zheng, Liming, et al.
Pubblicazione: (2024)
A Multimodal Interactive Framework for Science Assessment in the Era of Generative Artificial Intelligence
di: Yizhu Gao, et al.
Pubblicazione: (2025)
di: Yizhu Gao, et al.
Pubblicazione: (2025)
'No' Matters: Out-of-Distribution Detection in Multimodality Long Dialogue
di: Gao, Rena, et al.
Pubblicazione: (2024)
di: Gao, Rena, et al.
Pubblicazione: (2024)
Turing‐Structured Covalent Organic Framework Membranes for Fast and Precise Peptide Separations
di: Bingjie Gao, et al.
Pubblicazione: (2025)
di: Bingjie Gao, et al.
Pubblicazione: (2025)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
di: Jiang, Houcheng, et al.
Pubblicazione: (2026)
di: Jiang, Houcheng, et al.
Pubblicazione: (2026)
Understanding University Students' Use of Generative AI: The Roles of Demographics and Personality Traits
di: Deng, Newnew, et al.
Pubblicazione: (2025)
di: Deng, Newnew, et al.
Pubblicazione: (2025)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
di: Tang, Changli, et al.
Pubblicazione: (2026)
di: Tang, Changli, et al.
Pubblicazione: (2026)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
di: Lee, Yebin, et al.
Pubblicazione: (2024)
di: Lee, Yebin, et al.
Pubblicazione: (2024)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
di: Chai, Ying, et al.
Pubblicazione: (2025)
di: Chai, Ying, et al.
Pubblicazione: (2025)
MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in Dialogues
di: Guo, Diandian, et al.
Pubblicazione: (2026)
di: Guo, Diandian, et al.
Pubblicazione: (2026)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
di: Liu, Zeyu, et al.
Pubblicazione: (2026)
di: Liu, Zeyu, et al.
Pubblicazione: (2026)
Multimodal Trustworthy Semantic Communication for Audio-Visual Event Localization
di: Li, Yuandi, et al.
Pubblicazione: (2024)
di: Li, Yuandi, et al.
Pubblicazione: (2024)
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking
di: Pang, Renning, et al.
Pubblicazione: (2026)
di: Pang, Renning, et al.
Pubblicazione: (2026)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
di: Kang, Minjae, et al.
Pubblicazione: (2025)
di: Kang, Minjae, et al.
Pubblicazione: (2025)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction
di: Xu, Chao, et al.
Pubblicazione: (2026)
di: Xu, Chao, et al.
Pubblicazione: (2026)
Generating Adversarial Events: A Motion-Aware Point Cloud Framework
di: Ren, Hongwei, et al.
Pubblicazione: (2026)
di: Ren, Hongwei, et al.
Pubblicazione: (2026)
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
di: Kong, Zhe, et al.
Pubblicazione: (2025)
di: Kong, Zhe, et al.
Pubblicazione: (2025)
Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
di: Li, Zinuo, et al.
Pubblicazione: (2025)
di: Li, Zinuo, et al.
Pubblicazione: (2025)
AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
di: Chowdhury, Sanjoy, et al.
Pubblicazione: (2025)
di: Chowdhury, Sanjoy, et al.
Pubblicazione: (2025)
WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
di: Kang, Yongqi, et al.
Pubblicazione: (2025)
di: Kang, Yongqi, et al.
Pubblicazione: (2025)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
AudioX: A Unified Framework for Anything-to-Audio Generation
di: Tian, Zeyue, et al.
Pubblicazione: (2025)
di: Tian, Zeyue, et al.
Pubblicazione: (2025)
Polarization Management on Anisotropic Thin‐Film Lithium Niobate Platform
di: Kaixuan Chen, et al.
Pubblicazione: (2025)
di: Kaixuan Chen, et al.
Pubblicazione: (2025)
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
di: Gong, Kaixiong, et al.
Pubblicazione: (2024)
di: Gong, Kaixiong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
di: Deng, Xiang, et al.
Pubblicazione: (2026) -
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
di: Pang, Youxin, et al.
Pubblicazione: (2025) -
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
di: Shao, Ruizhi, et al.
Pubblicazione: (2024) -
Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
di: Deng, Xiang, et al.
Pubblicazione: (2024) -
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
di: Kong, Zhe, et al.
Pubblicazione: (2025)