Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Yuhang, Tang, Zixuan, Pang, Yunsheng, Ding, Beichen, Liu, Mengyuan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition
by: Wen, Yuhang, et al.
Published: (2024)
by: Wen, Yuhang, et al.
Published: (2024)
A Survey on Backbones for Deep Video Action Recognition
by: Tang, Zixuan, et al.
Published: (2024)
by: Tang, Zixuan, et al.
Published: (2024)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Improving Skeleton-based Action Recognition with Interactive Object Information
by: Wen, Hao, et al.
Published: (2025)
by: Wen, Hao, et al.
Published: (2025)
HDBN: A Novel Hybrid Dual-branch Network for Robust Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
by: Liu, Hongda, et al.
Published: (2025)
by: Liu, Hongda, et al.
Published: (2025)
HyLiFormer: Hyperbolic Linear Attention for Skeleton-based Human Action Recognition
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Skeleton-Based Human Action Recognition with Noisy Labels
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System
by: Karadag, Ozge Oztimur
Published: (2023)
by: Karadag, Ozge Oztimur
Published: (2023)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Including Semantic Information via Word Embeddings for Skeleton-based Action Recognition
by: Aganian, Dustin, et al.
Published: (2025)
by: Aganian, Dustin, et al.
Published: (2025)
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
by: Fang, Huang, et al.
Published: (2025)
by: Fang, Huang, et al.
Published: (2025)
Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
A Survey on 3D Skeleton-Based Action Recognition Using Learning Method
by: Ren, Bin, et al.
Published: (2020)
by: Ren, Bin, et al.
Published: (2020)
VG4D: Vision-Language Model Goes 4D Video Recognition
by: Deng, Zhichao, et al.
Published: (2024)
by: Deng, Zhichao, et al.
Published: (2024)
FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation
by: Zeng, Huajian, et al.
Published: (2026)
by: Zeng, Huajian, et al.
Published: (2026)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
by: Xie, Haozhe, et al.
Published: (2026)
by: Xie, Haozhe, et al.
Published: (2026)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
To Move or Not to Move: Constraint-based Planning Enables Zero-Shot Generalization for Interactive Navigation
by: Vashisth, Apoorva, et al.
Published: (2026)
by: Vashisth, Apoorva, et al.
Published: (2026)
Mixture of Horizons in Action Chunking
by: Jing, Dong, et al.
Published: (2025)
by: Jing, Dong, et al.
Published: (2025)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
by: Wang, Beichen, et al.
Published: (2024)
by: Wang, Beichen, et al.
Published: (2024)
OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
by: Ji, Yihong, et al.
Published: (2025)
by: Ji, Yihong, et al.
Published: (2025)
A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition
by: Yin, Ruoqi, et al.
Published: (2023)
by: Yin, Ruoqi, et al.
Published: (2023)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
by: Yuan, Haoran, et al.
Published: (2026)
by: Yuan, Haoran, et al.
Published: (2026)
Interactive Post-Training for Vision-Language-Action Models
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
3D Skeleton-Based Action Recognition: A Review
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
SSL-Interactions: Pretext Tasks for Interactive Trajectory Prediction
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026)
by: Maksutova, Aiza, et al.
Published: (2026)
Variational Contrastive Learning for Skeleton-based Action Recognition
by: Nguyen, Dang Dinh, et al.
Published: (2026)
by: Nguyen, Dang Dinh, et al.
Published: (2026)
SceneFoundry: Generating Interactive Infinite 3D Worlds
by: Chen, ChunTeng, et al.
Published: (2026)
by: Chen, ChunTeng, et al.
Published: (2026)
DiG-Net: Enhancing Human-Robot Interaction through Hyper-Range Dynamic Gesture Recognition in Assistive Robotics
by: Beeri, Eran Bamani, et al.
Published: (2025)
by: Beeri, Eran Bamani, et al.
Published: (2025)
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
by: Chen, Yandu, et al.
Published: (2025)
by: Chen, Yandu, et al.
Published: (2025)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
Explore Human Parsing Modality for Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
by: Zeng, Zichao, et al.
Published: (2026)
by: Zeng, Zichao, et al.
Published: (2026)
RALACs: Action Recognition in Autonomous Vehicles using Interaction Encoding and Optical Flow
by: Zhou, Eddy, et al.
Published: (2022)
by: Zhou, Eddy, et al.
Published: (2022)
Explicit Interaction for Fusion-Based Place Recognition
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
by: Gao, Chongkai, et al.
Published: (2025)
by: Gao, Chongkai, et al.
Published: (2025)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
by: Tang, Zuojin, et al.
Published: (2026)
by: Tang, Zuojin, et al.
Published: (2026)
Similar Items
-
CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition
by: Wen, Yuhang, et al.
Published: (2024) -
A Survey on Backbones for Deep Video Action Recognition
by: Tang, Zixuan, et al.
Published: (2024) -
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025) -
Improving Skeleton-based Action Recognition with Interactive Object Information
by: Wen, Hao, et al.
Published: (2025) -
HDBN: A Novel Hybrid Dual-branch Network for Robust Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)