MoChat: Joints-Grouped Spatio-Temporal Grounding LLM for Multi-Turn Motion Comprehension and Description
Fuente:
arXiv
Saved in:
| Main Authors: | Mo, Jiawei, Chen, Yixuan, Lin, Rifen, Ni, Yongkang, Zeng, Min, Hu, Xiping, Li, Min |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
by: Lin, Rifen, et al.
Published: (2025)
by: Lin, Rifen, et al.
Published: (2025)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
by: Gao, Hong, et al.
Published: (2025)
by: Gao, Hong, et al.
Published: (2025)
NeMo-map: Neural Implicit Flow Fields for Spatio-Temporal Motion Mapping
by: Zhu, Yufei, et al.
Published: (2025)
by: Zhu, Yufei, et al.
Published: (2025)
ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
Layer-wise Regularized Dropout for Neural Language Models
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
DALI: LLM-Agent Enhanced Dual-Stream Adaptive Leadership Identification for Group Recommendations
by: Song, Boxun, et al.
Published: (2026)
by: Song, Boxun, et al.
Published: (2026)
Multi-layer Motion Planning with Kinodynamic and Spatio-Temporal Constraints
by: Chatrola, Jeel, et al.
Published: (2025)
by: Chatrola, Jeel, et al.
Published: (2025)
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
by: Sun, Ye, et al.
Published: (2025)
by: Sun, Ye, et al.
Published: (2025)
MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling
by: Yuan, Weihao, et al.
Published: (2024)
by: Yuan, Weihao, et al.
Published: (2024)
HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance
by: Kong, Huanjun, et al.
Published: (2024)
by: Kong, Huanjun, et al.
Published: (2024)
Communication-Efficient Federated Learning by Exploiting Spatio-Temporal Correlations of Gradients
by: Zheng, Shenlong, et al.
Published: (2026)
by: Zheng, Shenlong, et al.
Published: (2026)
Forgetting before Learning: Utilizing Parametric Arithmetic for Knowledge Updating in Large Language Models
by: Ni, Shiwen, et al.
Published: (2023)
by: Ni, Shiwen, et al.
Published: (2023)
Spatio-Temporal Branching for Motion Prediction using Motion Increments
by: Wang, Jiexin, et al.
Published: (2023)
by: Wang, Jiexin, et al.
Published: (2023)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
by: Zhang, Yiqun, et al.
Published: (2026)
by: Zhang, Yiqun, et al.
Published: (2026)
Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor
by: Meng, Yapeng, et al.
Published: (2026)
by: Meng, Yapeng, et al.
Published: (2026)
Towards Gradient-based Time-Series Explanations through a SpatioTemporal Attention Network
by: Lee, Min Hun
Published: (2024)
by: Lee, Min Hun
Published: (2024)
Spatio-Temporal Motion Retargeting for Quadruped Robots
by: Yoon, Taerim, et al.
Published: (2024)
by: Yoon, Taerim, et al.
Published: (2024)
Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action Classification
by: Lin, Jimmy, et al.
Published: (2024)
by: Lin, Jimmy, et al.
Published: (2024)
Learning to Generalize Unseen Domains via Multi-Source Meta Learning for Text Classification
by: Hu, Yuxuan, et al.
Published: (2024)
by: Hu, Yuxuan, et al.
Published: (2024)
Spatio-Temporal Multi-Subgraph GCN for 3D Human Motion Prediction
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
by: Li, Xinhao, et al.
Published: (2025)
by: Li, Xinhao, et al.
Published: (2025)
Adaptive Spatio‐Temporal 3D Gaussian Splatting for Scenes with Oscillatory Motion
by: Petros Tzathas, et al.
Published: (2026)
by: Petros Tzathas, et al.
Published: (2026)
Towards Long-Form Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2026)
by: Gu, Xin, et al.
Published: (2026)
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
by: Zhao, Heng, et al.
Published: (2026)
by: Zhao, Heng, et al.
Published: (2026)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
LLM-Grounded Dynamic Task Planning with Hierarchical Temporal Logic for Human-Aware Multi-Robot Collaboration
by: Hu, Shuyuan, et al.
Published: (2026)
by: Hu, Shuyuan, et al.
Published: (2026)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
by: Huang, Wei-Jhe, et al.
Published: (2024)
by: Huang, Wei-Jhe, et al.
Published: (2024)
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
by: Liang, Qihua, et al.
Published: (2026)
by: Liang, Qihua, et al.
Published: (2026)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
Emotion Recognition from Skeleton Data: A Comprehensive Survey
by: Lu, Haifeng, et al.
Published: (2025)
by: Lu, Haifeng, et al.
Published: (2025)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
by: So, Yerim, et al.
Published: (2026)
by: So, Yerim, et al.
Published: (2026)
STAA: Spatio-Temporal Alignment Attention for Short-Term Precipitation Forecasting
by: Chen, Min, et al.
Published: (2024)
by: Chen, Min, et al.
Published: (2024)
Expression Syntax Information Bottleneck for Math Word Problems
by: Xiong, Jing, et al.
Published: (2023)
by: Xiong, Jing, et al.
Published: (2023)
Temporal Action Detection Model Compression by Progressive Block Drop
by: Chen, Xiaoyong, et al.
Published: (2025)
by: Chen, Xiaoyong, et al.
Published: (2025)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
by: Garg, Aaryan, et al.
Published: (2025)
by: Garg, Aaryan, et al.
Published: (2025)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
by: Pan, Ruihao, et al.
Published: (2026)
by: Pan, Ruihao, et al.
Published: (2026)
MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual Property
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Training on the Benchmark Is Not All You Need
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Similar Items
-
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
by: Lin, Rifen, et al.
Published: (2025) -
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
by: Gao, Hong, et al.
Published: (2025) -
NeMo-map: Neural Implicit Flow Fields for Spatio-Temporal Motion Mapping
by: Zhu, Yufei, et al.
Published: (2025) -
ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch
by: Chen, Jiawei, et al.
Published: (2025) -
Layer-wise Regularized Dropout for Neural Language Models
by: Ni, Shiwen, et al.
Published: (2024)