Temporal Visual Semantics-Induced Human Motion Understanding with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xing, Zheng, Zhao, Weibing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Large Motion Models with Million-Level Human Motions
von: Wang, Ye, et al.
Veröffentlicht: (2024)
von: Wang, Ye, et al.
Veröffentlicht: (2024)
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
von: Song, Jiafei, et al.
Veröffentlicht: (2026)
von: Song, Jiafei, et al.
Veröffentlicht: (2026)
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
MoFM: A Large-Scale Human Motion Foundation Model
von: Baharani, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Baharani, Mohammadreza, et al.
Veröffentlicht: (2025)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)
Spatio-Temporal Multi-Subgraph GCN for 3D Human Motion Prediction
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
Understanding the Effect of using Semantically Meaningful Tokens for Visual Representation Learning
von: Kalibhat, Neha, et al.
Veröffentlicht: (2024)
von: Kalibhat, Neha, et al.
Veröffentlicht: (2024)
Understanding Bias in Large-Scale Visual Datasets
von: Zeng, Boya, et al.
Veröffentlicht: (2024)
von: Zeng, Boya, et al.
Veröffentlicht: (2024)
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
von: Nguyen, Le Thien Phuc, et al.
Veröffentlicht: (2025)
von: Nguyen, Le Thien Phuc, et al.
Veröffentlicht: (2025)
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2026)
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2026)
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
von: Le, Anjie, et al.
Veröffentlicht: (2025)
von: Le, Anjie, et al.
Veröffentlicht: (2025)
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
von: Zhou, Wenhao, et al.
Veröffentlicht: (2025)
von: Zhou, Wenhao, et al.
Veröffentlicht: (2025)
Spatio-Temporal Branching for Motion Prediction using Motion Increments
von: Wang, Jiexin, et al.
Veröffentlicht: (2023)
von: Wang, Jiexin, et al.
Veröffentlicht: (2023)
Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models
von: Alam, Md Tanvirul
Veröffentlicht: (2026)
von: Alam, Md Tanvirul
Veröffentlicht: (2026)
Introducing Visual Perception Token into Multimodal Large Language Model
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
Visual Prompting in Multimodal Large Language Models: A Survey
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
Towards Understanding How Knowledge Evolves in Large Vision-Language Models
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
von: Wang, Sudong, et al.
Veröffentlicht: (2025)
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
von: Yin, Shaofeng, et al.
Veröffentlicht: (2025)
von: Yin, Shaofeng, et al.
Veröffentlicht: (2025)
Machine Learning Modeling for Multi-order Human Visual Motion Processing
von: Sun, Zitang, et al.
Veröffentlicht: (2025)
von: Sun, Zitang, et al.
Veröffentlicht: (2025)
Transformation of Biological Networks into Images via Semantic Cartography for Visual Interpretation and Scalable Deep Analysis
von: Mostafa, Sakib, et al.
Veröffentlicht: (2025)
von: Mostafa, Sakib, et al.
Veröffentlicht: (2025)
Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games
von: Krauss, Henrik, et al.
Veröffentlicht: (2026)
von: Krauss, Henrik, et al.
Veröffentlicht: (2026)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
von: Wang, Ruiyu, et al.
Veröffentlicht: (2025)
von: Wang, Ruiyu, et al.
Veröffentlicht: (2025)
Visual-Guided Key-Token Regularization for Multimodal Large Language Model Unlearning
von: Cai, Chengyi, et al.
Veröffentlicht: (2026)
von: Cai, Chengyi, et al.
Veröffentlicht: (2026)
Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models
von: Wang, Enguang, et al.
Veröffentlicht: (2026)
von: Wang, Enguang, et al.
Veröffentlicht: (2026)
LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation
von: Tang, Huadong, et al.
Veröffentlicht: (2024)
von: Tang, Huadong, et al.
Veröffentlicht: (2024)
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
von: Cao, Bin, et al.
Veröffentlicht: (2025)
von: Cao, Bin, et al.
Veröffentlicht: (2025)
Understanding Model Reprogramming for CLIP via Decoupling Visual Prompts
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
von: Fujitake, Masato
Veröffentlicht: (2024)
von: Fujitake, Masato
Veröffentlicht: (2024)
Enhancing Large Vision Model in Street Scene Semantic Understanding through Leveraging Posterior Optimization Trajectory
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2025)
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2025)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
von: Wang, Haoxiang, et al.
Veröffentlicht: (2023)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2023)
Fast-Slow Efficient Training for Multimodal Large Language Models via Visual Token Pruning
von: Zhang, Dingkun, et al.
Veröffentlicht: (2026)
von: Zhang, Dingkun, et al.
Veröffentlicht: (2026)
A Survey on Human Interaction Motion Generation
von: Sui, Kewei, et al.
Veröffentlicht: (2025)
von: Sui, Kewei, et al.
Veröffentlicht: (2025)
CoMusion: Towards Consistent Stochastic Human Motion Prediction via Motion Diffusion
von: Sun, Jiarui, et al.
Veröffentlicht: (2023)
von: Sun, Jiarui, et al.
Veröffentlicht: (2023)
Learning Visual-Semantic Subspace Representations
von: Moreira, Gabriel, et al.
Veröffentlicht: (2024)
von: Moreira, Gabriel, et al.
Veröffentlicht: (2024)
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
von: Zhou, Runjie, et al.
Veröffentlicht: (2026)
von: Zhou, Runjie, et al.
Veröffentlicht: (2026)
Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models
von: Mandelli, Lorenzo, et al.
Veröffentlicht: (2024)
von: Mandelli, Lorenzo, et al.
Veröffentlicht: (2024)
Understanding Task Transfer in Vision-Language Models
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
von: Sachdeva, Bhuvan, et al.
Veröffentlicht: (2025)
Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
von: Li, Sijie, et al.
Veröffentlicht: (2026)
von: Li, Sijie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scaling Large Motion Models with Million-Level Human Motions
von: Wang, Ye, et al.
Veröffentlicht: (2024) -
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
von: Song, Jiafei, et al.
Veröffentlicht: (2026) -
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2025) -
MoFM: A Large-Scale Human Motion Foundation Model
von: Baharani, Mohammadreza, et al.
Veröffentlicht: (2025) -
LinVT: Empower Your Image-level Large Language Model to Understand Videos
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)