LatentMove: Towards Complex Human Movement Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Taghipour, Ashkan, Ghahremani, Morteza, Bennamoun, Mohammed, Boussaid, Farid, Rekavandi, Aref Miri, Li, Zinuo, Ke, Qiuhong, Laga, Hamid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
by: Taghipour, Ashkan, et al.
Published: (2024)
by: Taghipour, Ashkan, et al.
Published: (2024)
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
by: Taghipour, Ashkan, et al.
Published: (2024)
by: Taghipour, Ashkan, et al.
Published: (2024)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026)
by: Taghipour, Ashkan, et al.
Published: (2026)
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
by: Taghipour, Ashkan, et al.
Published: (2025)
by: Taghipour, Ashkan, et al.
Published: (2025)
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
by: Jospin, Laurent Valentin, et al.
Published: (2021)
by: Jospin, Laurent Valentin, et al.
Published: (2021)
Towards Adaptive Subspace Detection in Heterogeneous Environment
by: Rekavandi, Aref Miri
Published: (2024)
by: Rekavandi, Aref Miri
Published: (2024)
Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis
by: Nizamani, Awais, et al.
Published: (2025)
by: Nizamani, Awais, et al.
Published: (2025)
Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
by: Li, Zinuo, et al.
Published: (2025)
by: Li, Zinuo, et al.
Published: (2025)
STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
by: Li, Zinuo, et al.
Published: (2026)
by: Li, Zinuo, et al.
Published: (2026)
A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-shaped Structures
by: Khanam, Tahmina, et al.
Published: (2024)
by: Khanam, Tahmina, et al.
Published: (2024)
3D Brain and Heart Volume Generative Models: A Survey
by: Liu, Yanbin, et al.
Published: (2022)
by: Liu, Yanbin, et al.
Published: (2022)
A Riemannian Framework for the Elastic Analysis of the Spatiotemporal Variability in the Shape and Structure of Tree-like 4D Objects
by: Khanam, Tahmina, et al.
Published: (2025)
by: Khanam, Tahmina, et al.
Published: (2025)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
by: Zhang, Xian, et al.
Published: (2025)
by: Zhang, Xian, et al.
Published: (2025)
DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition
by: Zhu, Jingmin, et al.
Published: (2025)
by: Zhu, Jingmin, et al.
Published: (2025)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
by: Rajapaksha, Uchitha, et al.
Published: (2024)
by: Rajapaksha, Uchitha, et al.
Published: (2024)
RS-Reg: Probabilistic and Robust Certified Regression Through Randomized Smoothing
by: Rekavandi, Aref Miri, et al.
Published: (2024)
by: Rekavandi, Aref Miri, et al.
Published: (2024)
Conditional Diffusion for 3D CT Volume Reconstruction from 2D X-rays
by: Rath, Martin, et al.
Published: (2026)
by: Rath, Martin, et al.
Published: (2026)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Hybrid Transformer-Mamba Architecture for Weakly Supervised Volumetric Medical Segmentation
by: Lyu, Yiheng, et al.
Published: (2025)
by: Lyu, Yiheng, et al.
Published: (2025)
Auxiliary Tasks Enhanced Dual-affinity Learning for Weakly Supervised Semantic Segmentation
by: Xu, Lian, et al.
Published: (2024)
by: Xu, Lian, et al.
Published: (2024)
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
by: Liu, Yanan, et al.
Published: (2025)
by: Liu, Yanan, et al.
Published: (2025)
Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection
by: Sagar, A S M Sharifuzzaman, et al.
Published: (2026)
by: Sagar, A S M Sharifuzzaman, et al.
Published: (2026)
SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
by: Wang, Ning, et al.
Published: (2026)
by: Wang, Ning, et al.
Published: (2026)
Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation
by: Zhu, Jingmin, et al.
Published: (2025)
by: Zhu, Jingmin, et al.
Published: (2025)
UIFormer: A Unified Transformer-based Framework for Incremental Few-Shot Object Detection and Instance Segmentation
by: Zhang, Chengyuan, et al.
Published: (2024)
by: Zhang, Chengyuan, et al.
Published: (2024)
Finite Element and Computational Fluid Dynamics Analysis of a Biodegradable Implant for Large Femoral Bone Defects
by: Sina Taghipour, et al.
Published: (2026)
by: Sina Taghipour, et al.
Published: (2026)
Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
by: Li, Yitong, et al.
Published: (2025)
by: Li, Yitong, et al.
Published: (2025)
Normal-guided Detail-Preserving Neural Implicit Function for High-Fidelity 3D Surface Reconstruction
by: Patel, Aarya, et al.
Published: (2024)
by: Patel, Aarya, et al.
Published: (2024)
CineLOG: A Training Free Approach for Cinematic Long Video Generation
by: Dehghanian, Zahra, et al.
Published: (2025)
by: Dehghanian, Zahra, et al.
Published: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
by: Dai, Yusheng, et al.
Published: (2026)
by: Dai, Yusheng, et al.
Published: (2026)
LongDiff: Training-Free Long Video Generation in One Go
by: Li, Zhuoling, et al.
Published: (2025)
by: Li, Zhuoling, et al.
Published: (2025)
The Golden Conjecture – A Fractal Approach to Predicting Prime Gaps / Die Goldene Vermutung – Ein fraktaler Ansatz zur Vorhersage von Primzahllücken
by: Aref, Mohammed
Published: (2025)
by: Aref, Mohammed
Published: (2025)
Latent Laplace Diffusion for Irregular Multivariate Time Series
by: You, Zinuo, et al.
Published: (2026)
by: You, Zinuo, et al.
Published: (2026)
Metric Unreliability in Multimodal Machine Unlearning: A Systematic Analysis and Principled Unified Score
by: Khan, Abdullah Ahmad, et al.
Published: (2026)
by: Khan, Abdullah Ahmad, et al.
Published: (2026)
Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation
by: Wang, Jiajun, et al.
Published: (2024)
by: Wang, Jiajun, et al.
Published: (2024)
Compiling a Model of Managers' Professional Meritocracy Based on Islamic and Iranian Teachings
by: Taghipour, Azar, et al.
Published: (2020)
by: Taghipour, Azar, et al.
Published: (2020)
Health‐promoting potentials of milk protein hydrolysates: A comparative analysis of camel, cow and donkey resources
by: Mohammad Javad Taghipour, et al.
Published: (2025)
by: Mohammad Javad Taghipour, et al.
Published: (2025)
DiaMond: Dementia Diagnosis with Multi-Modal Vision Transformers Using MRI and PET
by: Li, Yitong, et al.
Published: (2024)
by: Li, Yitong, et al.
Published: (2024)
Similar Items
-
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
by: Taghipour, Ashkan, et al.
Published: (2024) -
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
by: Taghipour, Ashkan, et al.
Published: (2024) -
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026) -
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
by: Taghipour, Ashkan, et al.
Published: (2025) -
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
by: Jospin, Laurent Valentin, et al.
Published: (2021)