GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Zhankai, Li, Bofan, Jin, Yukai, Li, Shuoqiu, Wang, Wei, Zhang, Yanfu, Gao, Shangqian, Liu, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
SMooGPT: Stylized Motion Generation using Large Language Models
by: Zhong, Lei, et al.
Published: (2025)
by: Zhong, Lei, et al.
Published: (2025)
AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
by: Han, Haonan, et al.
Published: (2024)
by: Han, Haonan, et al.
Published: (2024)
Scaling Large Motion Models with Million-Level Human Motions
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
FreeMotion: MoCap-Free Human Motion Synthesis with Multimodal Large Language Models
by: Zhang, Zhikai, et al.
Published: (2024)
by: Zhang, Zhikai, et al.
Published: (2024)
EventGPT: Event Stream Understanding with Multimodal Large Language Models
by: Liu, Shaoyu, et al.
Published: (2024)
by: Liu, Shaoyu, et al.
Published: (2024)
GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry
by: He, Xiankang, et al.
Published: (2026)
by: He, Xiankang, et al.
Published: (2026)
MotionPCM: Real-Time Motion Synthesis with Phased Consistency Model
by: Jiang, Lei, et al.
Published: (2025)
by: Jiang, Lei, et al.
Published: (2025)
MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning
by: Liu, Xiaoyang, et al.
Published: (2024)
by: Liu, Xiaoyang, et al.
Published: (2024)
MotionGPT3: Human Motion as a Second Modality
by: Zhu, Bingfan, et al.
Published: (2025)
by: Zhu, Bingfan, et al.
Published: (2025)
Large Motion Model for Unified Multi-Modal Motion Generation
by: Zhang, Mingyuan, et al.
Published: (2024)
by: Zhang, Mingyuan, et al.
Published: (2024)
Geometry-Aware Feature Matching for Large-Scale Structure from Motion
by: Chen, Gonglin, et al.
Published: (2024)
by: Chen, Gonglin, et al.
Published: (2024)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects
by: Shen, Licheng, et al.
Published: (2025)
by: Shen, Licheng, et al.
Published: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
by: Feng, Haoan, et al.
Published: (2026)
by: Feng, Haoan, et al.
Published: (2026)
Aligning Human Motion Generation with Human Perceptions
by: Wang, Haoru, et al.
Published: (2024)
by: Wang, Haoru, et al.
Published: (2024)
DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding
by: Zhang, Ning, et al.
Published: (2026)
by: Zhang, Ning, et al.
Published: (2026)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
by: Du, Yipeng, et al.
Published: (2025)
by: Du, Yipeng, et al.
Published: (2025)
EgoLM: Multi-Modal Language Model of Egocentric Motions
by: Hong, Fangzhou, et al.
Published: (2024)
by: Hong, Fangzhou, et al.
Published: (2024)
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
by: Lv, Jiaxi, et al.
Published: (2023)
by: Lv, Jiaxi, et al.
Published: (2023)
GeoDiffMM: Geometry-Guided Conditional Diffusion for Motion Magnification
by: Liu, Xuedeng, et al.
Published: (2025)
by: Liu, Xuedeng, et al.
Published: (2025)
Exploring Motion-Language Alignment for Text-driven Motion Generation
by: Gu, Ruxi, et al.
Published: (2026)
by: Gu, Ruxi, et al.
Published: (2026)
Temporal Visual Semantics-Induced Human Motion Understanding with Large Language Models
by: Xing, Zheng, et al.
Published: (2025)
by: Xing, Zheng, et al.
Published: (2025)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
by: Liu, Ruyang, et al.
Published: (2025)
by: Liu, Ruyang, et al.
Published: (2025)
Auto-Train-Once: Controller Network Guided Automatic Network Pruning from Scratch
by: Wu, Xidong, et al.
Published: (2024)
by: Wu, Xidong, et al.
Published: (2024)
MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
Restoration-Aligned Generative Flow Models for Blind Motion Deblurring
by: Kim, Insoo, et al.
Published: (2026)
by: Kim, Insoo, et al.
Published: (2026)
iMOVE: Instance-Motion-Aware Video Understanding
by: Li, Jiaze, et al.
Published: (2025)
by: Li, Jiaze, et al.
Published: (2025)
LLM-PCGC: Large Language Model-based Point Cloud Geometry Compression
by: Ye, Yuqi, et al.
Published: (2024)
by: Ye, Yuqi, et al.
Published: (2024)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
by: Zhang, Yaqi, et al.
Published: (2023)
by: Zhang, Yaqi, et al.
Published: (2023)
A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding
by: Wang, Yida, et al.
Published: (2025)
by: Wang, Yida, et al.
Published: (2025)
DreamMover: Leveraging the Prior of Diffusion Models for Image Interpolation with Large Motion
by: Shen, Liao, et al.
Published: (2024)
by: Shen, Liao, et al.
Published: (2024)
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
by: Ling, Xinran, et al.
Published: (2025)
by: Ling, Xinran, et al.
Published: (2025)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
by: Zhan, Yu-Wei, et al.
Published: (2025)
by: Zhan, Yu-Wei, et al.
Published: (2025)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation
by: Jin, Luoxu, et al.
Published: (2024)
by: Jin, Luoxu, et al.
Published: (2024)
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
by: Chen, Bohong, et al.
Published: (2025)
by: Chen, Bohong, et al.
Published: (2025)
IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
by: Li, Zewen, et al.
Published: (2025)
by: Li, Zewen, et al.
Published: (2025)
GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View Geometry
by: Le, Jiajun, et al.
Published: (2025)
by: Le, Jiajun, et al.
Published: (2025)
Similar Items
-
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
by: Wang, Yuan, et al.
Published: (2024) -
SMooGPT: Stylized Motion Generation using Large Language Models
by: Zhong, Lei, et al.
Published: (2025) -
AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
by: Han, Haonan, et al.
Published: (2024) -
Scaling Large Motion Models with Million-Level Human Motions
by: Wang, Ye, et al.
Published: (2024) -
FreeMotion: MoCap-Free Human Motion Synthesis with Multimodal Large Language Models
by: Zhang, Zhikai, et al.
Published: (2024)