Gespeichert in:
| Hauptverfasser: | Yu, Jiashuo, Yao, Yao, Chen, Boyu, Wang, Alex |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2606.01703 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation
von: Yao, Yao, et al.
Veröffentlicht: (2023)
von: Yao, Yao, et al.
Veröffentlicht: (2023)
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
von: Zhang, Haomin, et al.
Veröffentlicht: (2025)
von: Zhang, Haomin, et al.
Veröffentlicht: (2025)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
von: Ton, Tri, et al.
Veröffentlicht: (2025)
von: Ton, Tri, et al.
Veröffentlicht: (2025)
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2025)
von: Dai, Yusheng, et al.
Veröffentlicht: (2025)
Apollo: Unified Multi-Task Audio-Video Joint Generation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments
von: Ortigoso-Narro, Jorge, et al.
Veröffentlicht: (2025)
von: Ortigoso-Narro, Jorge, et al.
Veröffentlicht: (2025)
ST-GDance++: A Scalable Spatial-Temporal Diffusion for Long-Duration Group Choreography
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
von: Yang, Jialiang, et al.
Veröffentlicht: (2026)
Diffusion Models for Joint Audio-Video Generation
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
Do Joint Audio-Video Generation Models Understand Physics?
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
von: Tan, Weiting, et al.
Veröffentlicht: (2026)
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
von: Liu, Zhan, et al.
Veröffentlicht: (2026)
von: Liu, Zhan, et al.
Veröffentlicht: (2026)
DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided Distillation
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion
von: Ma, Jian, et al.
Veröffentlicht: (2024)
von: Ma, Jian, et al.
Veröffentlicht: (2024)
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
von: Jiang, Peiyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Peiyuan, et al.
Veröffentlicht: (2025)
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
360+x: A Panoptic Multi-modal Scene Understanding Dataset
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Learning Separable Hidden Unit Contributions for Speaker-Adaptive Lip-Reading
von: Luo, Songtao, et al.
Veröffentlicht: (2023)
von: Luo, Songtao, et al.
Veröffentlicht: (2023)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
von: Wang, Ke, et al.
Veröffentlicht: (2025)
von: Wang, Ke, et al.
Veröffentlicht: (2025)
Graph Propagated Projection Unlearning: A Unified Framework for Vision and Audio Discriminative Models
von: Pathak, Shreyansh, et al.
Veröffentlicht: (2026)
von: Pathak, Shreyansh, et al.
Veröffentlicht: (2026)
TRACE: Training-Free Partial Audio Deepfake Detection via Embedding Trajectory Analysis of Speech Foundation Models
von: Khan, Awais, et al.
Veröffentlicht: (2026)
von: Khan, Awais, et al.
Veröffentlicht: (2026)
Materialistic RIR: Material Conditioned Realistic RIR Generation
von: Saad, Mahnoor Fatima, et al.
Veröffentlicht: (2026)
von: Saad, Mahnoor Fatima, et al.
Veröffentlicht: (2026)
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
von: Khanal, Subash, et al.
Veröffentlicht: (2025)
von: Khanal, Subash, et al.
Veröffentlicht: (2025)
DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
von: Zaman, Sayeem Been, et al.
Veröffentlicht: (2025)
von: Zaman, Sayeem Been, et al.
Veröffentlicht: (2025)
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
von: Yang, Ziyue, et al.
Veröffentlicht: (2026)
von: Yang, Ziyue, et al.
Veröffentlicht: (2026)
IMSE: Efficient U-Net-based Speech Enhancement using Inception Depthwise Convolution and Amplitude-Aware Linear Attention
von: Tang, Xinxin, et al.
Veröffentlicht: (2025)
von: Tang, Xinxin, et al.
Veröffentlicht: (2025)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
von: Lokegaonkar, Vaibhavi, et al.
Veröffentlicht: (2026)
von: Lokegaonkar, Vaibhavi, et al.
Veröffentlicht: (2026)
Video-based Music Generation
von: Sulun, Serkan
Veröffentlicht: (2026)
von: Sulun, Serkan
Veröffentlicht: (2026)
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
von: Nie, Chang, et al.
Veröffentlicht: (2026)
von: Nie, Chang, et al.
Veröffentlicht: (2026)
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
von: Chen, Changan, et al.
Veröffentlicht: (2024)
von: Chen, Changan, et al.
Veröffentlicht: (2024)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
VoxAging: Continuously Tracking Speaker Aging with a Large-Scale Longitudinal Dataset in English and Mandarin
von: Ai, Zhiqi, et al.
Veröffentlicht: (2025)
von: Ai, Zhiqi, et al.
Veröffentlicht: (2025)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
von: Li, Qifei, et al.
Veröffentlicht: (2024)
von: Li, Qifei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation
von: Yao, Yao, et al.
Veröffentlicht: (2023) -
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
von: Zhang, Haomin, et al.
Veröffentlicht: (2025) -
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025) -
TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
von: Ton, Tri, et al.
Veröffentlicht: (2025) -
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2025)