MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Yipeng, Fan, Tiehan, Nan, Kepan, Xie, Rui, Zhou, Penghao, Li, Xiang, Yang, Jian, Yang, Zhenheng, Tai, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
by: Nan, Kepan, et al.
Published: (2024)
by: Nan, Kepan, et al.
Published: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025)
by: Song, Guorui, et al.
Published: (2025)
DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis
by: Chen, Peiyin, et al.
Published: (2025)
by: Chen, Peiyin, et al.
Published: (2025)
MotionCharacter: Fine-Grained Motion Controllable Human Video Generation
by: Fang, Haopeng, et al.
Published: (2024)
by: Fang, Haopeng, et al.
Published: (2024)
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
by: Tu, Chongjun, et al.
Published: (2025)
by: Tu, Chongjun, et al.
Published: (2025)
UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
by: Zhao, Chen, et al.
Published: (2025)
by: Zhao, Chen, et al.
Published: (2025)
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
by: Huang, Ziyuan, et al.
Published: (2024)
by: Huang, Ziyuan, et al.
Published: (2024)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
Show-o2: Improved Native Unified Multimodal Models
by: Xie, Jinheng, et al.
Published: (2025)
by: Xie, Jinheng, et al.
Published: (2025)
PartMotionEdit: Fine-Grained Text-Driven 3D Human Motion Editing via Part-Level Modulation
by: Yang, Yujie, et al.
Published: (2025)
by: Yang, Yujie, et al.
Published: (2025)
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
Motion Primitives Planning For Center-Articulated Vehicles
by: Hu, Jiangpeng, et al.
Published: (2024)
by: Hu, Jiangpeng, et al.
Published: (2024)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
Programmable Motion Generation for Open-Set Motion Control Tasks
by: Liu, Hanchao, et al.
Published: (2024)
by: Liu, Hanchao, et al.
Published: (2024)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
by: Li, Honglin, et al.
Published: (2024)
by: Li, Honglin, et al.
Published: (2024)
Trajectory Attention for Fine-grained Video Motion Control
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
Decoupling Contact for Fine-Grained Motion Style Transfer
by: Tang, Xiangjun, et al.
Published: (2024)
by: Tang, Xiangjun, et al.
Published: (2024)
Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation
by: Wang, Yin, et al.
Published: (2025)
by: Wang, Yin, et al.
Published: (2025)
FineXtrol: Controllable Motion Generation via Fine-Grained Text
by: Shen, Keming, et al.
Published: (2025)
by: Shen, Keming, et al.
Published: (2025)
PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
by: Wang, Penghao, et al.
Published: (2025)
by: Wang, Penghao, et al.
Published: (2025)
ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring
by: Lin, Xiaopeng, et al.
Published: (2025)
by: Lin, Xiaopeng, et al.
Published: (2025)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
by: Zhou, Wenrui, et al.
Published: (2025)
by: Zhou, Wenrui, et al.
Published: (2025)
MotionChain: Conversational Motion Controllers via Multimodal Prompts
by: Jiang, Biao, et al.
Published: (2024)
by: Jiang, Biao, et al.
Published: (2024)
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
by: Islam, Muhammad, et al.
Published: (2025)
by: Islam, Muhammad, et al.
Published: (2025)
BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion
by: Jia, Tianzhi, et al.
Published: (2026)
by: Jia, Tianzhi, et al.
Published: (2026)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Geometry-Guided Camera Motion Understanding in VideoLLMs
by: Feng, Haoan, et al.
Published: (2026)
by: Feng, Haoan, et al.
Published: (2026)
Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs
by: Yan, Qianqi, et al.
Published: (2025)
by: Yan, Qianqi, et al.
Published: (2025)
Improving Sparse IMU-based Motion Capture with Motion Label Smoothing
by: Meng, Zhaorui, et al.
Published: (2025)
by: Meng, Zhaorui, et al.
Published: (2025)
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
by: Lin, Boda, et al.
Published: (2026)
by: Lin, Boda, et al.
Published: (2026)
Understanding and Preserving Safety in Fine-Tuned LLMs
by: Zhang, Jiawen, et al.
Published: (2026)
by: Zhang, Jiawen, et al.
Published: (2026)
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework
by: Dong, Xin, et al.
Published: (2024)
by: Dong, Xin, et al.
Published: (2024)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
by: Zhu, Minghao, et al.
Published: (2023)
by: Zhu, Minghao, et al.
Published: (2023)
Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning
by: Chen, Hanmo, et al.
Published: (2026)
by: Chen, Hanmo, et al.
Published: (2026)
FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
Conductive Composite Hydrogel with Unsymmetrical Structure as Multimodal Triboelectric Nanogenerators for Machine Learning‐Assisted Motion
by: Yanqi Yin, et al.
Published: (2025)
by: Yanqi Yin, et al.
Published: (2025)
Similar Items
-
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
by: Nan, Kepan, et al.
Published: (2024) -
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024) -
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
by: Xie, Rui, et al.
Published: (2025) -
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025) -
DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis
by: Chen, Peiyin, et al.
Published: (2025)