Unified Video-Language Pre-training with Synchronized Audio
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mo, Shentong, Wang, Haofan, Li, Huaxia, Tang, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Text-to-Audio Generation Synchronized with Videos
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Semantic Grouping Network for Audio Source Separation
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Audio-visual Generalized Zero-shot Learning the Easy Way
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Continual Audio-Visual Sound Separation
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
Efficient Selective Audio Masked Multimodal Bottleneck Transformer for Audio-Video Classification
von: Zhu, Wentao
Veröffentlicht: (2024)
von: Zhu, Wentao
Veröffentlicht: (2024)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
Efficient Multiscale Multimodal Bottleneck Transformer for Audio-Video Classification
von: Zhu, Wentao
Veröffentlicht: (2024)
von: Zhu, Wentao
Veröffentlicht: (2024)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
AudioX: A Unified Framework for Anything-to-Audio Generation
von: Tian, Zeyue, et al.
Veröffentlicht: (2025)
von: Tian, Zeyue, et al.
Veröffentlicht: (2025)
AVTENet: A Human-Cognition-Inspired Audio-Visual Transformer-Based Ensemble Network for Video Deepfake Detection
von: Hashmi, Ammarah, et al.
Veröffentlicht: (2023)
von: Hashmi, Ammarah, et al.
Veröffentlicht: (2023)
Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
von: Zhou, Jinxing, et al.
Veröffentlicht: (2026)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2026)
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
von: Joo, Seohyun, et al.
Veröffentlicht: (2026)
von: Joo, Seohyun, et al.
Veröffentlicht: (2026)
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation
von: Su, Kun, et al.
Veröffentlicht: (2024)
von: Su, Kun, et al.
Veröffentlicht: (2024)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
von: Rai, Aashish, et al.
Veröffentlicht: (2024)
von: Rai, Aashish, et al.
Veröffentlicht: (2024)
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
Multimodal Sentiment Analysis based on Video and Audio Inputs
von: Fernandez, Antonio, et al.
Veröffentlicht: (2024)
von: Fernandez, Antonio, et al.
Veröffentlicht: (2024)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
von: Fang, Pengjun, et al.
Veröffentlicht: (2026)
von: Fang, Pengjun, et al.
Veröffentlicht: (2026)
3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control
von: Sha, Xuanmeng, et al.
Veröffentlicht: (2024)
von: Sha, Xuanmeng, et al.
Veröffentlicht: (2024)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling
von: Yang, Ruihan, et al.
Veröffentlicht: (2023)
von: Yang, Ruihan, et al.
Veröffentlicht: (2023)
Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
von: Lee, Junwon, et al.
Veröffentlicht: (2025)
von: Lee, Junwon, et al.
Veröffentlicht: (2025)
StereoSync: Spatially-Aware Stereo Audio Generation from Video
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2023)
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2023)
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
von: Hayakawa, Akio, et al.
Veröffentlicht: (2024)
von: Hayakawa, Akio, et al.
Veröffentlicht: (2024)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion
von: Sun, Yu, et al.
Veröffentlicht: (2025)
von: Sun, Yu, et al.
Veröffentlicht: (2025)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
von: Dai, Yusheng, et al.
Veröffentlicht: (2024)
von: Dai, Yusheng, et al.
Veröffentlicht: (2024)
Synchformer: Efficient Synchronization from Sparse Cues
von: Iashin, Vladimir, et al.
Veröffentlicht: (2024)
von: Iashin, Vladimir, et al.
Veröffentlicht: (2024)
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX
von: Xie, Liuyue, et al.
Veröffentlicht: (2025)
von: Xie, Liuyue, et al.
Veröffentlicht: (2025)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Text-to-Audio Generation Synchronized with Videos
von: Mo, Shentong, et al.
Veröffentlicht: (2024) -
Multi-scale Multi-instance Visual Sound Localization and Segmentation
von: Mo, Shentong, et al.
Veröffentlicht: (2024) -
Aligning Audio-Visual Joint Representations with an Agentic Workflow
von: Mo, Shentong, et al.
Veröffentlicht: (2024) -
Semantic Grouping Network for Audio Source Separation
von: Mo, Shentong, et al.
Veröffentlicht: (2024) -
Audio-visual Generalized Zero-shot Learning the Easy Way
von: Mo, Shentong, et al.
Veröffentlicht: (2024)