Universal Skeleton Understanding via Differentiable Rendering and MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ziyi, Li, Peiming, Wang, Xinshun, Tang, Yang, Ma, Kai-Kuang, Liu, Mengyuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
by: Wang, Xinshun, et al.
Published: (2026)
by: Wang, Xinshun, et al.
Published: (2026)
Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
by: Wang, Xinshun, et al.
Published: (2023)
by: Wang, Xinshun, et al.
Published: (2023)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition
by: Liu, Mengyuan, et al.
Published: (2026)
by: Liu, Mengyuan, et al.
Published: (2026)
ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion Models
by: Li, Peiming, et al.
Published: (2024)
by: Li, Peiming, et al.
Published: (2024)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Towards Universal Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026)
by: Kuang, Jidong, et al.
Published: (2026)
Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
by: Li, Peiming, et al.
Published: (2025)
by: Li, Peiming, et al.
Published: (2025)
Heterogeneous Skeleton-Based Action Representation Learning
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
A Simple Approach to Differentiable Rendering of SDFs
by: Wang, Zichen, et al.
Published: (2024)
by: Wang, Zichen, et al.
Published: (2024)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
MoPO: Incorporating Motion Prior for Occluded Human Mesh Recovery
by: Tang, Tao, et al.
Published: (2026)
by: Tang, Tao, et al.
Published: (2026)
Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026)
by: Kuang, Jidong, et al.
Published: (2026)
PFGS: High Fidelity Point Cloud Rendering via Feature Splatting
by: Wang, Jiaxu, et al.
Published: (2024)
by: Wang, Jiaxu, et al.
Published: (2024)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
by: Huang, Jincai, et al.
Published: (2026)
by: Huang, Jincai, et al.
Published: (2026)
HyLiFormer: Hyperbolic Linear Attention for Skeleton-based Human Action Recognition
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Understanding the Vulnerability of Skeleton-based Human Activity Recognition via Black-box Attack
by: Diao, Yunfeng, et al.
Published: (2022)
by: Diao, Yunfeng, et al.
Published: (2022)
A Survey on 3D Skeleton-Based Action Recognition Using Learning Method
by: Ren, Bin, et al.
Published: (2020)
by: Ren, Bin, et al.
Published: (2020)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
by: Sun, Yanpeng, et al.
Published: (2025)
by: Sun, Yanpeng, et al.
Published: (2025)
Zero-Shot Skeleton-based Action Recognition with Dual Visual-Text Alignment
by: Kuang, Jidong, et al.
Published: (2024)
by: Kuang, Jidong, et al.
Published: (2024)
Foundation Model for Skeleton-Based Human Action Understanding
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Affinity Contrastive Learning for Skeleton-based Human Activity Understanding
by: Liu, Hongda, et al.
Published: (2026)
by: Liu, Hongda, et al.
Published: (2026)
Gaussian Mesh Renderer for Lightweight Differentiable Rendering
by: Liu, Xinpeng, et al.
Published: (2026)
by: Liu, Xinpeng, et al.
Published: (2026)
Diff3DS: Generating View-Consistent 3D Sketch via Differentiable Curve Rendering
by: Zhang, Yibo, et al.
Published: (2024)
by: Zhang, Yibo, et al.
Published: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
by: Yang, Liu, et al.
Published: (2025)
by: Yang, Liu, et al.
Published: (2025)
SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
by: Liu, Hongda, et al.
Published: (2025)
by: Liu, Hongda, et al.
Published: (2025)
3D Skeleton-Based Action Recognition: A Review
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
by: Wang, Yonghui, et al.
Published: (2024)
by: Wang, Yonghui, et al.
Published: (2024)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
by: Zhu, Xiaorong, et al.
Published: (2025)
by: Zhu, Xiaorong, et al.
Published: (2025)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
by: Sun, Peiwen, et al.
Published: (2026)
by: Sun, Peiwen, et al.
Published: (2026)
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
by: Peng, Yibo, et al.
Published: (2026)
by: Peng, Yibo, et al.
Published: (2026)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
RepText: Rendering Visual Text via Replicating
by: Wang, Haofan, et al.
Published: (2025)
by: Wang, Haofan, et al.
Published: (2025)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
by: Wen, Yuhang, et al.
Published: (2023)
by: Wen, Yuhang, et al.
Published: (2023)
Similar Items
-
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
by: Wang, Xinshun, et al.
Published: (2026) -
Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
by: Wang, Xinshun, et al.
Published: (2023) -
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026) -
Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition
by: Liu, Mengyuan, et al.
Published: (2026) -
ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion Models
by: Li, Peiming, et al.
Published: (2024)