Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lin, Cai, Zefan, Zhou, Yufan, Mo, Shentong, Lin, Jinhong, Wu, Cheng-En, Wei, Yibing, Zhang, Yijing, Zhang, Ruiyi, Xiao, Wen, Sun, Tong, Hu, Junjie, Morgado, Pedro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audio-Synchronized Visual Animation
by: Zhang, Lin, et al.
Published: (2024)
by: Zhang, Lin, et al.
Published: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Audio-visual Generalized Zero-shot Learning the Easy Way
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Accelerating Augmentation Invariance Pretraining
by: Lin, Jinhong, et al.
Published: (2024)
by: Lin, Jinhong, et al.
Published: (2024)
Text-to-Audio Generation Synchronized with Videos
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training
by: Lin, Jinhong, et al.
Published: (2026)
by: Lin, Jinhong, et al.
Published: (2026)
Unified Video-Language Pre-training with Synchronized Audio
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Semantic Grouping Network for Audio Source Separation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing
by: Lin, Lixiang, et al.
Published: (2026)
by: Lin, Lixiang, et al.
Published: (2026)
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches
by: Wu, Cheng-En, et al.
Published: (2024)
by: Wu, Cheng-En, et al.
Published: (2024)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Towards Latent Masked Image Modeling for Self-Supervised Visual Representation Learning
by: Wei, Yibing, et al.
Published: (2024)
by: Wei, Yibing, et al.
Published: (2024)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
From Prototypes to General Distributions: An Efficient Curriculum for Masked Image Modeling
by: Lin, Jinhong, et al.
Published: (2024)
by: Lin, Jinhong, et al.
Published: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Continual Audio-Visual Sound Separation
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
by: Mo, Shentong
Published: (2024)
by: Mo, Shentong
Published: (2024)
Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning
by: Mo, Shentong
Published: (2026)
by: Mo, Shentong
Published: (2026)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
by: Mo, Shentong
Published: (2024)
by: Mo, Shentong
Published: (2024)
Customization Assistant for Text-to-image Generation
by: Zhou, Yufan, et al.
Published: (2023)
by: Zhou, Yufan, et al.
Published: (2023)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
A Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation
by: Fernandes, Jose Geraldo, et al.
Published: (2024)
by: Fernandes, Jose Geraldo, et al.
Published: (2024)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
GMAIL: Generative Modality Alignment for generated Image Learning
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
SaDiT: Efficient Protein Backbone Design via Latent Structural Tokenization and Diffusion Transformers
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
by: Pendyala, Varsha, et al.
Published: (2025)
by: Pendyala, Varsha, et al.
Published: (2025)
Delta Attention Residuals
by: Luo, Cheng, et al.
Published: (2026)
by: Luo, Cheng, et al.
Published: (2026)
Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
by: Ma, Xingpei, et al.
Published: (2025)
by: Ma, Xingpei, et al.
Published: (2025)
A Large-scale Medical Visual Task Adaptation Benchmark
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
by: Xu, Zhihua, et al.
Published: (2025)
by: Xu, Zhihua, et al.
Published: (2025)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Similar Items
-
Audio-Synchronized Visual Animation
by: Zhang, Lin, et al.
Published: (2024) -
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024) -
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
by: Mo, Shentong, et al.
Published: (2026) -
Audio-visual Generalized Zero-shot Learning the Easy Way
by: Mo, Shentong, et al.
Published: (2024) -
Accelerating Augmentation Invariance Pretraining
by: Lin, Jinhong, et al.
Published: (2024)