Audio-Synchronized Visual Animation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lin, Mo, Shentong, Zhang, Yijing, Morgado, Pedro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Audio-visual Generalized Zero-shot Learning the Easy Way
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Text-to-Audio Generation Synchronized with Videos
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
Unified Video-Language Pre-training with Synchronized Audio
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Semantic Grouping Network for Audio Source Separation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
by: Mo, Shentong
Published: (2024)
by: Mo, Shentong
Published: (2024)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing
by: Lin, Lixiang, et al.
Published: (2026)
by: Lin, Lixiang, et al.
Published: (2026)
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
by: Xu, Mingwang, et al.
Published: (2024)
by: Xu, Mingwang, et al.
Published: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
by: Xu, Zhihua, et al.
Published: (2025)
by: Xu, Zhihua, et al.
Published: (2025)
Continual Audio-Visual Sound Separation
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
by: Mo, Shentong
Published: (2024)
by: Mo, Shentong
Published: (2024)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
A Large-scale Medical Visual Task Adaptation Benchmark
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
A Synchronized Audio-Visual Multi-View Capture System
by: Shi, Xiangwei, et al.
Published: (2026)
by: Shi, Xiangwei, et al.
Published: (2026)
ALIVE: Animate Your World with Lifelike Audio-Video Generation
by: Guo, Ying, et al.
Published: (2026)
by: Guo, Ying, et al.
Published: (2026)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation
by: Meng, Dechao, et al.
Published: (2025)
by: Meng, Dechao, et al.
Published: (2025)
LayerAnimate: Layer-level Control for Animation
by: Yang, Yuxue, et al.
Published: (2025)
by: Yang, Yuxue, et al.
Published: (2025)
LinguaLinker: Audio-Driven Portraits Animation with Implicit Facial Control Enhancement
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
by: Huang, Donglin, et al.
Published: (2025)
by: Huang, Donglin, et al.
Published: (2025)
GMAIL: Generative Modality Alignment for generated Image Learning
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
EmoFace: Audio-driven Emotional 3D Face Animation
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
AnimateAnywhere: Rouse the Background in Human Image Animation
by: Liu, Xiaoyu, et al.
Published: (2025)
by: Liu, Xiaoyu, et al.
Published: (2025)
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
by: Zhen, Dingcheng, et al.
Published: (2025)
by: Zhen, Dingcheng, et al.
Published: (2025)
IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation
by: Xu, Zhufeng, et al.
Published: (2026)
by: Xu, Zhufeng, et al.
Published: (2026)
Text-Animator: Controllable Visual Text Video Generation
by: Liu, Lin, et al.
Published: (2024)
by: Liu, Lin, et al.
Published: (2024)
Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Towards Latent Masked Image Modeling for Self-Supervised Visual Representation Learning
by: Wei, Yibing, et al.
Published: (2024)
by: Wei, Yibing, et al.
Published: (2024)
Similar Items
-
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
by: Zhang, Lin, et al.
Published: (2025) -
Audio-visual Generalized Zero-shot Learning the Easy Way
by: Mo, Shentong, et al.
Published: (2024) -
Text-to-Audio Generation Synchronized with Videos
by: Mo, Shentong, et al.
Published: (2024) -
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
by: Mo, Shentong, et al.
Published: (2026) -
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)