Human Motion Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Lei, Jia, Sen, Wang, Jianhao, Jiang, Zhongyu, Zhou, Feng, Dai, Ju, Zhang, Tianfang, Wu, Zongkai, Hwang, Jenq-Neng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bayesian Optimization for Controlled Image Editing via LLMs
by: Cai, Chengkun, et al.
Published: (2025)
by: Cai, Chengkun, et al.
Published: (2025)
Graph Canvas for Controllable 3D Scene Generation
by: Liu, Libin, et al.
Published: (2024)
by: Liu, Libin, et al.
Published: (2024)
RAM: Recover Any 3D Human Motion in-the-Wild
by: Jia, Sen, et al.
Published: (2026)
by: Jia, Sen, et al.
Published: (2026)
ScalingGaussian: Enhancing 3D Content Creation with Generative Gaussian Splatting
by: Chen, Shen, et al.
Published: (2024)
by: Chen, Shen, et al.
Published: (2024)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
by: Jiang, Zhongyu, et al.
Published: (2025)
by: Jiang, Zhongyu, et al.
Published: (2025)
From Global to Local: Rethinking CLIP Feature Aggregation for Person Re-Identification
by: Zheng, Aotian, et al.
Published: (2026)
by: Zheng, Aotian, et al.
Published: (2026)
Learning to Learn Weight Generation via Local Consistency Diffusion
by: Guan, Yunchuan, et al.
Published: (2025)
by: Guan, Yunchuan, et al.
Published: (2025)
UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning
by: Jiang, Zhongyu, et al.
Published: (2025)
by: Jiang, Zhongyu, et al.
Published: (2025)
RT-Pose: A 4D Radar Tensor-based 3D Human Pose Estimation and Localization Benchmark
by: Ho, Yuan-Hao, et al.
Published: (2024)
by: Ho, Yuan-Hao, et al.
Published: (2024)
MambaMOT: State-Space Model as Motion Predictor for Multi-Object Tracking
by: Huang, Hsiang-Wei, et al.
Published: (2024)
by: Huang, Hsiang-Wei, et al.
Published: (2024)
SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory
by: Yang, Cheng-Yen, et al.
Published: (2024)
by: Yang, Cheng-Yen, et al.
Published: (2024)
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
by: Li, Lei, et al.
Published: (2025)
by: Li, Lei, et al.
Published: (2025)
The Role of Deductive and Inductive Reasoning in Large Language Models
by: Cai, Chengkun, et al.
Published: (2024)
by: Cai, Chengkun, et al.
Published: (2024)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
by: Zhang, Tianfang, et al.
Published: (2024)
by: Zhang, Tianfang, et al.
Published: (2024)
Reasoning Matters for 3D Visual Grounding
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
MPM: A Unified 2D-3D Human Pose Representation via Masked Pose Modeling
by: Zhang, Zhenyu, et al.
Published: (2023)
by: Zhang, Zhenyu, et al.
Published: (2023)
Tree Counting by Bridging 3D Point Clouds with Imagery
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
EVAN: Evolutional Video Streaming Adaptation via Neural Representation
by: Liu, Mufan, et al.
Published: (2024)
by: Liu, Mufan, et al.
Published: (2024)
GTA: Global Tracklet Association for Multi-Object Tracking in Sports
by: Sun, Jiacheng, et al.
Published: (2024)
by: Sun, Jiacheng, et al.
Published: (2024)
StickMotion: Generating 3D Human Motions by Drawing a Stickman
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
by: Wu, Song, et al.
Published: (2026)
by: Wu, Song, et al.
Published: (2026)
Position-aware Guided Point Cloud Completion with CLIP Model
by: Zhou, Feng, et al.
Published: (2024)
by: Zhou, Feng, et al.
Published: (2024)
Recent Advances in Embedding Methods for Multi-Object Tracking: A Survey
by: Wang, Gaoang, et al.
Published: (2022)
by: Wang, Gaoang, et al.
Published: (2022)
A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Video
by: Yang, Cheng-Yen, et al.
Published: (2024)
by: Yang, Cheng-Yen, et al.
Published: (2024)
Human-Object Interaction from Human-Level Instructions
by: Wu, Zhen, et al.
Published: (2024)
by: Wu, Zhen, et al.
Published: (2024)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
by: Wu, Changti, et al.
Published: (2026)
by: Wu, Changti, et al.
Published: (2026)
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
by: Xie, Hongxia, et al.
Published: (2024)
by: Xie, Hongxia, et al.
Published: (2024)
DAug: Diffusion-based Channel Augmentation for Radiology Image Retrieval and Classification
by: Jin, Ying, et al.
Published: (2024)
by: Jin, Ying, et al.
Published: (2024)
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026)
by: Dong, Mingkang, et al.
Published: (2026)
Warehouse Spatial Question Answering with LLM Agent
by: Huang, Hsiang-Wei, et al.
Published: (2025)
by: Huang, Hsiang-Wei, et al.
Published: (2025)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Beyond Human-prompting: Adaptive Prompt Tuning with Semantic Alignment for Anomaly Detection
by: Chen, Pi-Wei, et al.
Published: (2025)
by: Chen, Pi-Wei, et al.
Published: (2025)
Causal Reasoning Elicits Controllable 3D Scene Generation
by: Chen, Shen, et al.
Published: (2025)
by: Chen, Shen, et al.
Published: (2025)
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
by: Du, Lingxiao, et al.
Published: (2025)
by: Du, Lingxiao, et al.
Published: (2025)
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
by: Liu, Jinlin, et al.
Published: (2024)
by: Liu, Jinlin, et al.
Published: (2024)
PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology
by: Wu, Xiaomin, et al.
Published: (2024)
by: Wu, Xiaomin, et al.
Published: (2024)
MANTIS: Interleaved Multi-Image Instruction Tuning
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
P2P-Insole: Human Pose Estimation Using Foot Pressure Distribution and Motion Sensors
by: Watanabe, Atsuya, et al.
Published: (2025)
by: Watanabe, Atsuya, et al.
Published: (2025)
Similar Items
-
Bayesian Optimization for Controlled Image Editing via LLMs
by: Cai, Chengkun, et al.
Published: (2025) -
Graph Canvas for Controllable 3D Scene Generation
by: Liu, Libin, et al.
Published: (2024) -
RAM: Recover Any 3D Human Motion in-the-Wild
by: Jia, Sen, et al.
Published: (2026) -
ScalingGaussian: Enhancing 3D Content Creation with Generative Gaussian Splatting
by: Chen, Shen, et al.
Published: (2024) -
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
by: Jiang, Zhongyu, et al.
Published: (2025)