SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pham, Kien T., He, Yingqing, Xing, Yazhou, Chen, Qifeng, Chen, Long |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
von: Fang, Pengjun, et al.
Veröffentlicht: (2026)
von: Fang, Pengjun, et al.
Veröffentlicht: (2026)
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
von: Pina, Leonardo, et al.
Veröffentlicht: (2024)
von: Pina, Leonardo, et al.
Veröffentlicht: (2024)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
von: Zhang, Fan, et al.
Veröffentlicht: (2023)
von: Zhang, Fan, et al.
Veröffentlicht: (2023)
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
von: Bhattacharya, Aneesh, et al.
Veröffentlicht: (2023)
von: Bhattacharya, Aneesh, et al.
Veröffentlicht: (2023)
MusicScore: A Dataset for Music Score Modeling and Generation
von: Lin, Yuheng, et al.
Veröffentlicht: (2024)
von: Lin, Yuheng, et al.
Veröffentlicht: (2024)
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
von: Jung, Jongmin, et al.
Veröffentlicht: (2025)
von: Jung, Jongmin, et al.
Veröffentlicht: (2025)
Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
von: Xie, Siyi, et al.
Veröffentlicht: (2025)
von: Xie, Siyi, et al.
Veröffentlicht: (2025)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
von: Ji, Xiaozhong, et al.
Veröffentlicht: (2024)
von: Ji, Xiaozhong, et al.
Veröffentlicht: (2024)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
von: Ren, Yong, et al.
Veröffentlicht: (2024)
von: Ren, Yong, et al.
Veröffentlicht: (2024)
Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
von: Han, Zhen, et al.
Veröffentlicht: (2025)
von: Han, Zhen, et al.
Veröffentlicht: (2025)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
von: Lei, Ke, et al.
Veröffentlicht: (2026)
von: Lei, Ke, et al.
Veröffentlicht: (2026)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module
von: Wu, Kam Man, et al.
Veröffentlicht: (2025)
von: Wu, Kam Man, et al.
Veröffentlicht: (2025)
Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video Deepfake Dataset
von: Kaur, Sukhandeep, et al.
Veröffentlicht: (2024)
von: Kaur, Sukhandeep, et al.
Veröffentlicht: (2024)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation
von: Chan, Nolan, et al.
Veröffentlicht: (2026)
von: Chan, Nolan, et al.
Veröffentlicht: (2026)
Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
von: Chen, Gehui, et al.
Veröffentlicht: (2024)
von: Chen, Gehui, et al.
Veröffentlicht: (2024)
StereoFoley: Object-Aware Stereo Audio Generation from Video
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
von: Enge, Kajetan, et al.
Veröffentlicht: (2024)
von: Enge, Kajetan, et al.
Veröffentlicht: (2024)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2024)
LoVA: Long-form Video-to-Audio Generation
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025)
von: Gu, Ke, et al.
Veröffentlicht: (2025)
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
von: Ren, Yong, et al.
Veröffentlicht: (2025)
von: Ren, Yong, et al.
Veröffentlicht: (2025)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024)
von: Lin, Jinwei
Veröffentlicht: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
von: Xing, Yazhou, et al.
Veröffentlicht: (2024) -
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
von: Fang, Pengjun, et al.
Veröffentlicht: (2026) -
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
von: Pina, Leonardo, et al.
Veröffentlicht: (2024) -
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
von: Zhang, Fan, et al.
Veröffentlicht: (2023) -
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)