JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Xiaohu, Zhou, Hao, Yang, Qiangpeng, Wen, Shilei, Han, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
by: Huang, Xiaohu, et al.
Published: (2024)
by: Huang, Xiaohu, et al.
Published: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
by: He, Haoyang, et al.
Published: (2025)
by: He, Haoyang, et al.
Published: (2025)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
by: Liu, Kai, et al.
Published: (2026)
by: Liu, Kai, et al.
Published: (2026)
JoPano: Unified Panorama Generation via Joint Modeling
by: Feng, Wancheng, et al.
Published: (2025)
by: Feng, Wancheng, et al.
Published: (2025)
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
by: Li, Ruibin, et al.
Published: (2025)
by: Li, Ruibin, et al.
Published: (2025)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition
by: Fritsch, Stefan Gerd, et al.
Published: (2024)
by: Fritsch, Stefan Gerd, et al.
Published: (2024)
Identity as Presence: Towards Appearance and Voice Personalized Joint Audio-Video Generation
by: Chen, Yingjie, et al.
Published: (2026)
by: Chen, Yingjie, et al.
Published: (2026)
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
by: Tu, Shuyuan, et al.
Published: (2026)
by: Tu, Shuyuan, et al.
Published: (2026)
JoIN: Joint GANs Inversion for Intrinsic Image Decomposition
by: Shah, Viraj, et al.
Published: (2023)
by: Shah, Viraj, et al.
Published: (2023)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
by: Ma, Bingqi, et al.
Published: (2026)
by: Ma, Bingqi, et al.
Published: (2026)
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
by: Wang, Duomin, et al.
Published: (2025)
by: Wang, Duomin, et al.
Published: (2025)
Apollo: Unified Multi-Task Audio-Video Joint Generation
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
by: Xiong, Junwen, et al.
Published: (2024)
by: Xiong, Junwen, et al.
Published: (2024)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
by: Huang, Xiaohu, et al.
Published: (2024)
by: Huang, Xiaohu, et al.
Published: (2024)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
by: Han, Gyuwon, et al.
Published: (2026)
by: Han, Gyuwon, et al.
Published: (2026)
Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective
by: Zhu, Duowang, et al.
Published: (2025)
by: Zhu, Duowang, et al.
Published: (2025)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning
by: Cheng, Xin, et al.
Published: (2026)
by: Cheng, Xin, et al.
Published: (2026)
JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image Enhancement
by: Wu, Yuhui, et al.
Published: (2023)
by: Wu, Yuhui, et al.
Published: (2023)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
by: Liu, Jialun, et al.
Published: (2026)
by: Liu, Jialun, et al.
Published: (2026)
Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
by: Liu, Xiaohong, et al.
Published: (2025)
by: Liu, Xiaohong, et al.
Published: (2025)
MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection
by: Hu, Mengxue, et al.
Published: (2025)
by: Hu, Mengxue, et al.
Published: (2025)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
by: Li, Ruibin, et al.
Published: (2026)
by: Li, Ruibin, et al.
Published: (2026)
DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
by: Hogue, Steven, et al.
Published: (2024)
by: Hogue, Steven, et al.
Published: (2024)
JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
by: Zheng, Haojie, et al.
Published: (2026)
by: Zheng, Haojie, et al.
Published: (2026)
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
by: Cao, Zhe, et al.
Published: (2025)
by: Cao, Zhe, et al.
Published: (2025)
ProJo4D: Progressive Joint Optimization for Sparse-View Inverse Physics Estimation
by: Rho, Daniel, et al.
Published: (2025)
by: Rho, Daniel, et al.
Published: (2025)
OmniCam: Unified Multimodal Video Generation via Camera Control
by: Yang, Xiaoda, et al.
Published: (2025)
by: Yang, Xiaoda, et al.
Published: (2025)
JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
by: Son, Taein, et al.
Published: (2024)
by: Son, Taein, et al.
Published: (2024)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Similar Items
-
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026) -
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
by: Huang, Xiaohu, et al.
Published: (2024) -
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026) -
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
by: He, Haoyang, et al.
Published: (2025) -
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
by: Liu, Kai, et al.
Published: (2026)