Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
Fuente:
arXiv
Guardado en:
| Autores principales: | Taghipour, Ashkan, Ghahremani, Morteza, Li, Zinuo, Laga, Hamid, Boussaid, Farid, Bennamoun, Mohammed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LatentMove: Towards Complex Human Movement Video Generation
por: Taghipour, Ashkan, et al.
Publicado: (2025)
por: Taghipour, Ashkan, et al.
Publicado: (2025)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
por: Taghipour, Ashkan, et al.
Publicado: (2024)
por: Taghipour, Ashkan, et al.
Publicado: (2024)
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
por: Taghipour, Ashkan, et al.
Publicado: (2024)
por: Taghipour, Ashkan, et al.
Publicado: (2024)
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
por: Taghipour, Ashkan, et al.
Publicado: (2025)
por: Taghipour, Ashkan, et al.
Publicado: (2025)
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
por: Jospin, Laurent Valentin, et al.
Publicado: (2021)
por: Jospin, Laurent Valentin, et al.
Publicado: (2021)
Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis
por: Nizamani, Awais, et al.
Publicado: (2025)
por: Nizamani, Awais, et al.
Publicado: (2025)
Human Motion Video Generation: A Survey
por: Xue, Haiwei, et al.
Publicado: (2025)
por: Xue, Haiwei, et al.
Publicado: (2025)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
por: Chai, Zenghao, et al.
Publicado: (2024)
por: Chai, Zenghao, et al.
Publicado: (2024)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
por: Zhang, Zhongwei, et al.
Publicado: (2025)
por: Zhang, Zhongwei, et al.
Publicado: (2025)
Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion
por: Wang, Xinghan, et al.
Publicado: (2024)
por: Wang, Xinghan, et al.
Publicado: (2024)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
por: Guan, Jiazhi, et al.
Publicado: (2025)
por: Guan, Jiazhi, et al.
Publicado: (2025)
A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-shaped Structures
por: Khanam, Tahmina, et al.
Publicado: (2024)
por: Khanam, Tahmina, et al.
Publicado: (2024)
SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
por: Wang, Ning, et al.
Publicado: (2026)
por: Wang, Ning, et al.
Publicado: (2026)
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
por: Zhao, Sihan, et al.
Publicado: (2025)
por: Zhao, Sihan, et al.
Publicado: (2025)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
por: Zhang, Xian, et al.
Publicado: (2025)
por: Zhang, Xian, et al.
Publicado: (2025)
SOSControl: Enhancing Human Motion Generation through Saliency-Aware Symbolic Orientation and Timing Control
por: Au, Ho Yin, et al.
Publicado: (2025)
por: Au, Ho Yin, et al.
Publicado: (2025)
TAVGBench: Benchmarking Text to Audible-Video Generation
por: Mao, Yuxin, et al.
Publicado: (2024)
por: Mao, Yuxin, et al.
Publicado: (2024)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
por: Moradi, Morteza, et al.
Publicado: (2024)
por: Moradi, Morteza, et al.
Publicado: (2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
por: Zhang, Yuang, et al.
Publicado: (2024)
por: Zhang, Yuang, et al.
Publicado: (2024)
Generating Attribute-Aware Human Motions from Textual Prompt
por: Wang, Xinghan, et al.
Publicado: (2025)
por: Wang, Xinghan, et al.
Publicado: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
por: Gao, Jiayi, et al.
Publicado: (2025)
por: Gao, Jiayi, et al.
Publicado: (2025)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
por: Yuan, Shenghai, et al.
Publicado: (2024)
por: Yuan, Shenghai, et al.
Publicado: (2024)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
por: Chen, Liyang, et al.
Publicado: (2025)
por: Chen, Liyang, et al.
Publicado: (2025)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
por: Cheng, Shihao, et al.
Publicado: (2026)
por: Cheng, Shihao, et al.
Publicado: (2026)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
por: Zhang, Beiyuan, et al.
Publicado: (2024)
por: Zhang, Beiyuan, et al.
Publicado: (2024)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
por: Razlighi, AmirHossein Naghi, et al.
Publicado: (2026)
por: Razlighi, AmirHossein Naghi, et al.
Publicado: (2026)
Scene-Text Grounding for Text-Based Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2024)
por: Zhou, Sheng, et al.
Publicado: (2024)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
por: Xie, Zequn, et al.
Publicado: (2026)
por: Xie, Zequn, et al.
Publicado: (2026)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
por: Chen, Weifeng, et al.
Publicado: (2023)
por: Chen, Weifeng, et al.
Publicado: (2023)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
por: Shi, Haoyuan, et al.
Publicado: (2026)
por: Shi, Haoyuan, et al.
Publicado: (2026)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
por: Ilaslan, Muhammet Furkan, et al.
Publicado: (2024)
por: Ilaslan, Muhammet Furkan, et al.
Publicado: (2024)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
por: Zhang, Zongye, et al.
Publicado: (2025)
por: Zhang, Zongye, et al.
Publicado: (2025)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023)
por: Wang, Zhouxia, et al.
Publicado: (2023)
Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions
por: Dehghanian, Zahra, et al.
Publicado: (2025)
por: Dehghanian, Zahra, et al.
Publicado: (2025)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
por: Wu, Xun, et al.
Publicado: (2024)
por: Wu, Xun, et al.
Publicado: (2024)
PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
por: Jin, Chuhao, et al.
Publicado: (2025)
por: Jin, Chuhao, et al.
Publicado: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2025)
por: Zhou, Sheng, et al.
Publicado: (2025)
ViMo: Generating Motions from Casual Videos
por: Qiu, Liangdong, et al.
Publicado: (2024)
por: Qiu, Liangdong, et al.
Publicado: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
por: Ling, Jun, et al.
Publicado: (2024)
por: Ling, Jun, et al.
Publicado: (2024)
Ejemplares similares
-
LatentMove: Towards Complex Human Movement Video Generation
por: Taghipour, Ashkan, et al.
Publicado: (2025) -
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
por: Taghipour, Ashkan, et al.
Publicado: (2024) -
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
por: Taghipour, Ashkan, et al.
Publicado: (2024) -
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
por: Taghipour, Ashkan, et al.
Publicado: (2025) -
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
por: Jospin, Laurent Valentin, et al.
Publicado: (2021)