Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Hao, Shao, Ling, Zhang, Zhenyu, Van Gool, Luc, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
Towards Online Real-Time Memory-based Video Inpainting Transformers
by: Thiry, Guillaume, et al.
Published: (2024)
by: Thiry, Guillaume, et al.
Published: (2024)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
by: Ma, Qi, et al.
Published: (2025)
by: Ma, Qi, et al.
Published: (2025)
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
by: Yang, Ziyue, et al.
Published: (2026)
by: Yang, Ziyue, et al.
Published: (2026)
Any Image Restoration via Efficient Spatial-Frequency Degradation Adaptation
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
MatIR: A Hybrid Mamba-Transformer Image Restoration Model
by: Wen, Juan, et al.
Published: (2025)
by: Wen, Juan, et al.
Published: (2025)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Key-Graph Transformer for Image Restoration
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
ShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
by: Yang, Kaixing, et al.
Published: (2025)
by: Yang, Kaixing, et al.
Published: (2025)
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
by: Sun, Guolei, et al.
Published: (2022)
by: Sun, Guolei, et al.
Published: (2022)
Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Appearance Graphs
by: Segu, Mattia, et al.
Published: (2024)
by: Segu, Mattia, et al.
Published: (2024)
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
by: Dey, Sombit, et al.
Published: (2024)
by: Dey, Sombit, et al.
Published: (2024)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
by: Gao, Shida, et al.
Published: (2025)
by: Gao, Shida, et al.
Published: (2025)
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
by: He, Yikang, et al.
Published: (2026)
by: He, Yikang, et al.
Published: (2026)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
by: Mahdi, Mohammad, et al.
Published: (2025)
by: Mahdi, Mohammad, et al.
Published: (2025)
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
by: Rigo, Andrea, et al.
Published: (2025)
by: Rigo, Andrea, et al.
Published: (2025)
HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud
by: Cheng, Wencan, et al.
Published: (2024)
by: Cheng, Wencan, et al.
Published: (2024)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025)
by: Zuo, Zhi, et al.
Published: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
by: Sacilotti, André, et al.
Published: (2024)
by: Sacilotti, André, et al.
Published: (2024)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
by: Motamed, Saman, et al.
Published: (2024)
by: Motamed, Saman, et al.
Published: (2024)
Video Depth Propagation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
X-Dancer: Expressive Music to Human Dance Video Generation
by: Chen, Zeyuan, et al.
Published: (2025)
by: Chen, Zeyuan, et al.
Published: (2025)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Reverse Personalization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
Similar Items
-
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024) -
Enhanced Multi-Scale Cross-Attention for Person Image Generation
by: Tang, Hao, et al.
Published: (2025) -
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
by: Nguyen, Quang, et al.
Published: (2025) -
Towards Online Real-Time Memory-based Video Inpainting Transformers
by: Thiry, Guillaume, et al.
Published: (2024) -
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)