Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Feizhen, Wu, Yu, Lin, Yutian, Du, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
von: S, Sridhar, et al.
Veröffentlicht: (2025)
von: S, Sridhar, et al.
Veröffentlicht: (2025)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
Do Joint Audio-Video Generation Models Understand Physics?
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
Diffusion Models for Joint Audio-Video Generation
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Apollo: Unified Multi-Task Audio-Video Joint Generation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Audio-visual Event Localization on Portrait Mode Short Videos
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
Style-Preserving Lip Sync via Audio-Aware Style Reference
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Interactive Video Generation via Domain Adaptation
von: Rawal, Ishaan, et al.
Veröffentlicht: (2025)
von: Rawal, Ishaan, et al.
Veröffentlicht: (2025)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
ReCorD: Reasoning and Correcting Diffusion for HOI Generation
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
von: Zhou, Pengyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Pengyuan, et al.
Veröffentlicht: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024)
von: Yan, Xin, et al.
Veröffentlicht: (2024)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024)
von: Gu, Jing, et al.
Veröffentlicht: (2024)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
von: Qing, Yuan, et al.
Veröffentlicht: (2026)
von: Qing, Yuan, et al.
Veröffentlicht: (2026)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2026)
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2026)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
von: S, Sridhar, et al.
Veröffentlicht: (2025) -
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026) -
Do Joint Audio-Video Generation Models Understand Physics?
von: Cui, Zijun, et al.
Veröffentlicht: (2026) -
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025) -
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025)