Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Xinghan, Kang, Zixi, Mu, Yadong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generating Attribute-Aware Human Motions from Textual Prompt
por: Wang, Xinghan, et al.
Publicado: (2025)
por: Wang, Xinghan, et al.
Publicado: (2025)
Biomechanics-Guided Residual Approach to Generalizable Human Motion Generation and Estimation
por: Kang, Zixi, et al.
Publicado: (2025)
por: Kang, Zixi, et al.
Publicado: (2025)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
por: Taghipour, Ashkan, et al.
Publicado: (2026)
por: Taghipour, Ashkan, et al.
Publicado: (2026)
PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
por: Jin, Chuhao, et al.
Publicado: (2025)
por: Jin, Chuhao, et al.
Publicado: (2025)
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
por: Zhao, Sihan, et al.
Publicado: (2025)
por: Zhao, Sihan, et al.
Publicado: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2024)
por: Zhou, Sheng, et al.
Publicado: (2024)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
por: Li, Wenrui, et al.
Publicado: (2024)
por: Li, Wenrui, et al.
Publicado: (2024)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
por: Razlighi, AmirHossein Naghi, et al.
Publicado: (2026)
por: Razlighi, AmirHossein Naghi, et al.
Publicado: (2026)
Human Motion Video Generation: A Survey
por: Xue, Haiwei, et al.
Publicado: (2025)
por: Xue, Haiwei, et al.
Publicado: (2025)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
por: Yang, Danni, et al.
Publicado: (2024)
por: Yang, Danni, et al.
Publicado: (2024)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
por: Messina, Nicola, et al.
Publicado: (2024)
por: Messina, Nicola, et al.
Publicado: (2024)
InstructHumans: Editing Animated 3D Human Textures with Instructions
por: Zhu, Jiayin, et al.
Publicado: (2024)
por: Zhu, Jiayin, et al.
Publicado: (2024)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
por: Chai, Zenghao, et al.
Publicado: (2024)
por: Chai, Zenghao, et al.
Publicado: (2024)
BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion
por: Jia, Tianzhi, et al.
Publicado: (2026)
por: Jia, Tianzhi, et al.
Publicado: (2026)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
por: Zhang, Zongye, et al.
Publicado: (2025)
por: Zhang, Zongye, et al.
Publicado: (2025)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
por: Wang, Yilin, et al.
Publicado: (2025)
por: Wang, Yilin, et al.
Publicado: (2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
por: Zhang, Zhongwei, et al.
Publicado: (2025)
por: Zhang, Zhongwei, et al.
Publicado: (2025)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
por: Zhang, Yuang, et al.
Publicado: (2024)
por: Zhang, Yuang, et al.
Publicado: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
por: Ling, Jun, et al.
Publicado: (2024)
por: Ling, Jun, et al.
Publicado: (2024)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
por: Ilaslan, Muhammet Furkan, et al.
Publicado: (2024)
por: Ilaslan, Muhammet Furkan, et al.
Publicado: (2024)
SOSControl: Enhancing Human Motion Generation through Saliency-Aware Symbolic Orientation and Timing Control
por: Au, Ho Yin, et al.
Publicado: (2025)
por: Au, Ho Yin, et al.
Publicado: (2025)
MorphText: Deep Morphology Regularized Arbitrary-shape Scene Text Detection
por: Xu, Chengpei, et al.
Publicado: (2024)
por: Xu, Chengpei, et al.
Publicado: (2024)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
por: Chen, Yaru, et al.
Publicado: (2025)
por: Chen, Yaru, et al.
Publicado: (2025)
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
por: Li, Yuzhen, et al.
Publicado: (2025)
por: Li, Yuzhen, et al.
Publicado: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
por: Zhou, Sheng, et al.
Publicado: (2025)
por: Zhou, Sheng, et al.
Publicado: (2025)
LocoMotion: Learning Motion-Focused Video-Language Representations
por: Doughty, Hazel, et al.
Publicado: (2024)
por: Doughty, Hazel, et al.
Publicado: (2024)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
por: Yang, Shuyu, et al.
Publicado: (2024)
por: Yang, Shuyu, et al.
Publicado: (2024)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
por: Wu, Xun, et al.
Publicado: (2024)
por: Wu, Xun, et al.
Publicado: (2024)
Multiscale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines
por: Liu, Junle, et al.
Publicado: (2025)
por: Liu, Junle, et al.
Publicado: (2025)
Deep Compositional Phase Diffusion for Long Motion Sequence Generation
por: Au, Ho Yin, et al.
Publicado: (2025)
por: Au, Ho Yin, et al.
Publicado: (2025)
Boosting Temporal Sentence Grounding via Causal Inference
por: Tang, Kefan, et al.
Publicado: (2025)
por: Tang, Kefan, et al.
Publicado: (2025)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
por: Pramanick, Shraman, et al.
Publicado: (2025)
por: Pramanick, Shraman, et al.
Publicado: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
por: Zhang, Zhenxing, et al.
Publicado: (2024)
por: Zhang, Zhenxing, et al.
Publicado: (2024)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
por: Huang, Yiheng, et al.
Publicado: (2024)
por: Huang, Yiheng, et al.
Publicado: (2024)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
por: Guan, Runwei, et al.
Publicado: (2024)
por: Guan, Runwei, et al.
Publicado: (2024)
MambaMOS: LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space Model
por: Zeng, Kang, et al.
Publicado: (2024)
por: Zeng, Kang, et al.
Publicado: (2024)
L-LBVC: Long-Term Motion Estimation and Prediction for Learned Bi-Directional Video Compression
por: Zhai, Yongqi, et al.
Publicado: (2025)
por: Zhai, Yongqi, et al.
Publicado: (2025)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
por: Wang, Yuhao, et al.
Publicado: (2024)
por: Wang, Yuhao, et al.
Publicado: (2024)
Seeing Text in the Dark: Algorithm and Benchmark
por: Xu, Chengpei, et al.
Publicado: (2024)
por: Xu, Chengpei, et al.
Publicado: (2024)
Ejemplares similares
-
Generating Attribute-Aware Human Motions from Textual Prompt
por: Wang, Xinghan, et al.
Publicado: (2025) -
Biomechanics-Guided Residual Approach to Generalizable Human Motion Generation and Estimation
por: Kang, Zixi, et al.
Publicado: (2025) -
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
por: Taghipour, Ashkan, et al.
Publicado: (2026) -
PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
por: Jin, Chuhao, et al.
Publicado: (2025) -
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
por: Zhao, Sihan, et al.
Publicado: (2025)