Let Your Video Listen to Your Music!
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xinyu, Gong, Dong, Duan, Zicheng, Hengel, Anton van den, Liu, Lingqiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
by: Zhang, Jiaxu, et al.
Published: (2025)
by: Zhang, Jiaxu, et al.
Published: (2025)
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
by: Zhang, Beiyuan, et al.
Published: (2024)
by: Zhang, Beiyuan, et al.
Published: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
MVBIND: Self-Supervised Music Recommendation For Videos Via Embedding Space Binding
by: Teng, Jiajie, et al.
Published: (2024)
by: Teng, Jiajie, et al.
Published: (2024)
DiffBrush:Just Painting the Art by Your Hands
by: Chu, Jiaming, et al.
Published: (2025)
by: Chu, Jiaming, et al.
Published: (2025)
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Gorgeous: Create Your Desired Character Facial Makeup from Any Ideas
by: Sii, Jia Wei, et al.
Published: (2024)
by: Sii, Jia Wei, et al.
Published: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
by: Gao, Lishuai, et al.
Published: (2024)
by: Gao, Lishuai, et al.
Published: (2024)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
by: Liang, Yingshan, et al.
Published: (2025)
by: Liang, Yingshan, et al.
Published: (2025)
POINTS: Improving Your Vision-language Model with Affordable Strategies
by: Liu, Yuan, et al.
Published: (2024)
by: Liu, Yuan, et al.
Published: (2024)
Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
by: Kim, Geewook, et al.
Published: (2025)
by: Kim, Geewook, et al.
Published: (2025)
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
by: Cong, Gaoxiang, et al.
Published: (2025)
by: Cong, Gaoxiang, et al.
Published: (2025)
Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis
by: Liu, Miao, et al.
Published: (2026)
by: Liu, Miao, et al.
Published: (2026)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
by: Meng, Jiahao, et al.
Published: (2026)
by: Meng, Jiahao, et al.
Published: (2026)
CPSL: Representing Volumetric Video via Content-Promoted Scene Layers
by: Hu, Kaiyuan, et al.
Published: (2025)
by: Hu, Kaiyuan, et al.
Published: (2025)
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
by: Zhu, Xilei, et al.
Published: (2024)
by: Zhu, Xilei, et al.
Published: (2024)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
by: Yin, Zhiyu, et al.
Published: (2026)
by: Yin, Zhiyu, et al.
Published: (2026)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
by: Shi, Haoyuan, et al.
Published: (2026)
by: Shi, Haoyuan, et al.
Published: (2026)
Adaptive 3D Gaussian Splatting Video Streaming
by: Gong, Han, et al.
Published: (2025)
by: Gong, Han, et al.
Published: (2025)
Vlogger: Make Your Dream A Vlog
by: Zhuang, Shaobin, et al.
Published: (2024)
by: Zhuang, Shaobin, et al.
Published: (2024)
Advanced Learning-Based Inter Prediction for Future Video Coding
by: Zhao, Yanchen, et al.
Published: (2024)
by: Zhao, Yanchen, et al.
Published: (2024)
Towards Explainable Partial-AIGC Image Quality Assessment
by: Qian, Jiaying, et al.
Published: (2025)
by: Qian, Jiaying, et al.
Published: (2025)
PRVR: Partially Relevant Video Retrieval
by: Chen, Xianke, et al.
Published: (2022)
by: Chen, Xianke, et al.
Published: (2022)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
by: Zhang, Long, et al.
Published: (2025)
by: Zhang, Long, et al.
Published: (2025)
VCoME: Verbal Video Composition with Multimodal Editing Effects
by: Gong, Weibo, et al.
Published: (2024)
by: Gong, Weibo, et al.
Published: (2024)
AIS 2024 Challenge on Video Quality Assessment of User-Generated Content: Methods and Results
by: Conde, Marcos V., et al.
Published: (2024)
by: Conde, Marcos V., et al.
Published: (2024)
Consistency-aware Fake Videos Detection on Short Video Platforms
by: Wang, Junxi, et al.
Published: (2025)
by: Wang, Junxi, et al.
Published: (2025)
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
by: Li, Dong, et al.
Published: (2025)
by: Li, Dong, et al.
Published: (2025)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
by: Dai, Guangyu, et al.
Published: (2025)
by: Dai, Guangyu, et al.
Published: (2025)
Failures to Surface Harmful Contents in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2025)
by: Cao, Yuxin, et al.
Published: (2025)
Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
by: Li, Yuze, et al.
Published: (2026)
by: Li, Yuze, et al.
Published: (2026)
G-Refine: A General Quality Refiner for Text-to-Image Generation
by: Li, Chunyi, et al.
Published: (2024)
by: Li, Chunyi, et al.
Published: (2024)
Region-Constraint In-Context Generation for Instructional Video Editing
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
Casual3DHDR: Deblurring High Dynamic Range 3D Gaussian Splatting from Casually Captured Videos
by: Gong, Shucheng, et al.
Published: (2025)
by: Gong, Shucheng, et al.
Published: (2025)
Similar Items
-
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026) -
Bridging Your Imagination with Audio-Video Generation via a Unified Director
by: Zhang, Jiaxu, et al.
Published: (2025) -
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
by: Zhang, Xinyu, et al.
Published: (2025) -
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
by: Zhang, Beiyuan, et al.
Published: (2024) -
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)