STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Xiaoyu, Huang, Po-Yao, Liang, Junwei, de Melo, Celso M., Hauptmann, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarially Masked Video Consistency for Unsupervised Domain Adaptation
by: Zhu, Xiaoyu, et al.
Published: (2024)
by: Zhu, Xiaoyu, et al.
Published: (2024)
MoCap-to-Visual Domain Adaptation for Efficient Human Mesh Estimation from 2D Keypoints
by: Uguz, Bedirhan, et al.
Published: (2024)
by: Uguz, Bedirhan, et al.
Published: (2024)
MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors
by: Zhang, He, et al.
Published: (2024)
by: Zhang, He, et al.
Published: (2024)
MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond
by: Ren, Shenghao, et al.
Published: (2025)
by: Ren, Shenghao, et al.
Published: (2025)
FreeMotion: MoCap-Free Human Motion Synthesis with Multimodal Large Language Models
by: Zhang, Zhikai, et al.
Published: (2024)
by: Zhang, Zhikai, et al.
Published: (2024)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
by: Bouzid, Hamza, et al.
Published: (2023)
by: Bouzid, Hamza, et al.
Published: (2023)
No MoCap Needed: Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual Prompts
by: Macaluso, Girolamo, et al.
Published: (2025)
by: Macaluso, Girolamo, et al.
Published: (2025)
Mesquite MoCap: Democratizing Real-Time Motion Capture with Affordable, Bodyworn IoT Sensors and WebXR SLAM
by: Vanani, Poojan, et al.
Published: (2025)
by: Vanani, Poojan, et al.
Published: (2025)
DancingBox: A Lightweight MoCap System for Character Animation from Physical Proxies
by: Yuan, Haocheng, et al.
Published: (2026)
by: Yuan, Haocheng, et al.
Published: (2026)
Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
by: Zhu, Xiaoyu, et al.
Published: (2024)
by: Zhu, Xiaoyu, et al.
Published: (2024)
Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition
by: Qian, Zefeng, et al.
Published: (2025)
by: Qian, Zefeng, et al.
Published: (2025)
SNN-Driven Multimodal Human Action Recognition via Sparse Spatial-Temporal Data Fusion
by: Zheng, Naichuan, et al.
Published: (2025)
by: Zheng, Naichuan, et al.
Published: (2025)
When Spatial meets Temporal in Action Recognition
by: Chen, Huilin, et al.
Published: (2024)
by: Chen, Huilin, et al.
Published: (2024)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Spatial-Temporal Perception with Causal Inference for Naturalistic Driving Action Recognition
by: Chang, Qing, et al.
Published: (2025)
by: Chang, Qing, et al.
Published: (2025)
Human Action Recognition (HAR) Using Skeleton-based Spatial Temporal Relative Transformer Network: ST-RTR
by: Mehmood, Faisal, et al.
Published: (2024)
by: Mehmood, Faisal, et al.
Published: (2024)
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
SHARDeg: A Benchmark for Skeletal Human Action Recognition in Degraded Scenarios
by: Malzard, Simon, et al.
Published: (2025)
by: Malzard, Simon, et al.
Published: (2025)
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
by: Zhou, Jiaming, et al.
Published: (2023)
by: Zhou, Jiaming, et al.
Published: (2023)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
by: Sabathier, Remy, et al.
Published: (2026)
by: Sabathier, Remy, et al.
Published: (2026)
UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?
by: Xie, Jianyang, et al.
Published: (2025)
by: Xie, Jianyang, et al.
Published: (2025)
T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
by: Zhang, Shijie, et al.
Published: (2024)
by: Zhang, Shijie, et al.
Published: (2024)
Stack Transformer Based Spatial-Temporal Attention Model for Dynamic Sign Language and Fingerspelling Recognition
by: Hirooka, Koki, et al.
Published: (2025)
by: Hirooka, Koki, et al.
Published: (2025)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
LORTSAR: Low-Rank Transformer for Skeleton-based Action Recognition
by: Oraki, Soroush, et al.
Published: (2024)
by: Oraki, Soroush, et al.
Published: (2024)
Unsupervised Spatial-Temporal Feature Enrichment and Fidelity Preservation Network for Skeleton based Action Recognition
by: Li, Chuankun, et al.
Published: (2024)
by: Li, Chuankun, et al.
Published: (2024)
Ultron: Enabling Temporal Geometry Compression of 3D Mesh Sequences using Temporal Correspondence and Mesh Deformation
by: Zhu, Haichao
Published: (2024)
by: Zhu, Haichao
Published: (2024)
IIP-Transformer: Intra-Inter-Part Transformer for Skeleton-Based Action Recognition
by: Wang, Qingtian, et al.
Published: (2021)
by: Wang, Qingtian, et al.
Published: (2021)
Spatial Hierarchy and Temporal Attention Guided Cross Masking for Self-supervised Skeleton-based Action Recognition
by: Yin, Xinpeng, et al.
Published: (2024)
by: Yin, Xinpeng, et al.
Published: (2024)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
Synthetic-to-Real Domain Adaptation for Action Recognition: A Dataset and Baseline Performances
by: Reddy, Arun V., et al.
Published: (2023)
by: Reddy, Arun V., et al.
Published: (2023)
CorrNet+: Sign Language Recognition and Translation via Spatial-Temporal Correlation
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
by: Qiu, Yicheng, et al.
Published: (2026)
by: Qiu, Yicheng, et al.
Published: (2026)
One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
by: Yang, Siyuan, et al.
Published: (2023)
by: Yang, Siyuan, et al.
Published: (2023)
MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons
by: Gong, Kehong, et al.
Published: (2026)
by: Gong, Kehong, et al.
Published: (2026)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
AM Flow: Adapters for Temporal Processing in Action Recognition
by: Agrawal, Tanay, et al.
Published: (2024)
by: Agrawal, Tanay, et al.
Published: (2024)
Similar Items
-
Adversarially Masked Video Consistency for Unsupervised Domain Adaptation
by: Zhu, Xiaoyu, et al.
Published: (2024) -
MoCap-to-Visual Domain Adaptation for Efficient Human Mesh Estimation from 2D Keypoints
by: Uguz, Bedirhan, et al.
Published: (2024) -
MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors
by: Zhang, He, et al.
Published: (2024) -
MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond
by: Ren, Shenghao, et al.
Published: (2025) -
FreeMotion: MoCap-Free Human Motion Synthesis with Multimodal Large Language Models
by: Zhang, Zhikai, et al.
Published: (2024)