Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Teng, Yu, Zhentao, Zhang, Guozhen, Su, Zihan, Zhou, Zhengguang, Zhang, Youliang, Zhou, Yuan, Lu, Qinglin, Yi, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
by: Zhang, Youliang, et al.
Published: (2026)
by: Zhang, Youliang, et al.
Published: (2026)
OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
by: Liang, Sen, et al.
Published: (2025)
by: Liang, Sen, et al.
Published: (2025)
StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
by: Sun, Zhiyao, et al.
Published: (2025)
by: Sun, Zhiyao, et al.
Published: (2025)
UltraGen: High-Resolution Video Generation with Hierarchical Attention
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Arbitrary Generative Video Interpolation
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
by: Liang, Sen, et al.
Published: (2026)
by: Liang, Sen, et al.
Published: (2026)
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)
by: Wu, Xiaofei, et al.
Published: (2025)
ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
by: Yi, Ran, et al.
Published: (2025)
by: Yi, Ran, et al.
Published: (2025)
PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
by: Wang, Ruiyan, et al.
Published: (2025)
by: Wang, Ruiyan, et al.
Published: (2025)
Video Generation Models Are Good Latent Reward Models
by: Mi, Xiaoyue, et al.
Published: (2025)
by: Mi, Xiaoyue, et al.
Published: (2025)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
HarmonyDream: Task Harmonization Inside World Models
by: Ma, Haoyu, et al.
Published: (2023)
by: Ma, Haoyu, et al.
Published: (2023)
Evolution of Video Generative Foundations
by: Hu, Teng, et al.
Published: (2026)
by: Hu, Teng, et al.
Published: (2026)
Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
by: Fan, Ziqing, et al.
Published: (2024)
by: Fan, Ziqing, et al.
Published: (2024)
DiffHarmony: Latent Diffusion Model Meets Image Harmonization
by: Zhou, Pengfei, et al.
Published: (2024)
by: Zhou, Pengfei, et al.
Published: (2024)
USV: Unified Sparsification for Accelerating Video Diffusion Models
by: Wu, Xinjian, et al.
Published: (2025)
by: Wu, Xinjian, et al.
Published: (2025)
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
by: Hong, Fa-Ting, et al.
Published: (2025)
by: Hong, Fa-Ting, et al.
Published: (2025)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
by: Ma, Bingqi, et al.
Published: (2026)
by: Ma, Bingqi, et al.
Published: (2026)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
InstanceV: Instance-Level Video Generation
by: Chen, Yuheng, et al.
Published: (2025)
by: Chen, Yuheng, et al.
Published: (2025)
Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
by: Zeng, Jianshu, et al.
Published: (2025)
by: Zeng, Jianshu, et al.
Published: (2025)
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
by: Zhou, Zitang, et al.
Published: (2025)
by: Zhou, Zitang, et al.
Published: (2025)
Non-interference analysis of bounded labeled Petri nets
by: Ran, Ning, et al.
Published: (2025)
by: Ran, Ning, et al.
Published: (2025)
Neuromorphic Synergy for Video Binarization
by: Lin, Shijie, et al.
Published: (2024)
by: Lin, Shijie, et al.
Published: (2024)
AWCP: A Workspace Delegation Protocol for Deep-Engagement Collaboration across Remote Agents
by: Nie, Xiaohang, et al.
Published: (2026)
by: Nie, Xiaohang, et al.
Published: (2026)
Self‐Supervised Image Harmonization via Region‐Aware Harmony Classification
by: Chenyang Tian, et al.
Published: (2025)
by: Chenyang Tian, et al.
Published: (2025)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
by: Lian, Jiesong, et al.
Published: (2025)
by: Lian, Jiesong, et al.
Published: (2025)
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
by: Xie, Tianyidan, et al.
Published: (2026)
by: Xie, Tianyidan, et al.
Published: (2026)
TaoCache: Structure-Maintained Video Generation Acceleration
by: Fan, Zhentao, et al.
Published: (2025)
by: Fan, Zhentao, et al.
Published: (2025)
HarmonicAttack: An Adaptive Cross-Domain Audio Watermark Removal
by: Li, Kexin, et al.
Published: (2025)
by: Li, Kexin, et al.
Published: (2025)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
by: Zhang, Youliang, et al.
Published: (2025)
by: Zhang, Youliang, et al.
Published: (2025)
HarmoniAD: Harmonizing Local Structures and Global Semantics for Anomaly Detection
by: Zhang, Naiqi, et al.
Published: (2026)
by: Zhang, Naiqi, et al.
Published: (2026)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
by: Zhou, S. Z., et al.
Published: (2025)
by: Zhou, S. Z., et al.
Published: (2025)
Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web
by: Nie, Xiaohang, et al.
Published: (2026)
by: Nie, Xiaohang, et al.
Published: (2026)
ALIVE: Animate Your World with Lifelike Audio-Video Generation
by: Guo, Ying, et al.
Published: (2026)
by: Guo, Ying, et al.
Published: (2026)
Similar Items
-
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
by: Hu, Teng, et al.
Published: (2025) -
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025) -
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
by: Hu, Teng, et al.
Published: (2025) -
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
by: Zhang, Youliang, et al.
Published: (2026) -
OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
by: Liang, Sen, et al.
Published: (2025)