MALT: Multi-scale Action Learning Transformer for Online Action Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhipeng, Wang, Ruoyu, Tan, Yang, Xie, Liping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors
by: Yuan, Kaishen, et al.
Published: (2024)
by: Yuan, Kaishen, et al.
Published: (2024)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
by: Xie, Liping, et al.
Published: (2025)
by: Xie, Liping, et al.
Published: (2025)
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
by: Liu, Chenyv, et al.
Published: (2026)
by: Liu, Chenyv, et al.
Published: (2026)
Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic
by: Tang, Jianwei, et al.
Published: (2025)
by: Tang, Jianwei, et al.
Published: (2025)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
by: Yang, Zhuoyi, et al.
Published: (2026)
by: Yang, Zhuoyi, et al.
Published: (2026)
Multivariate Gaussian Representation Learning for Medical Action Evaluation
by: Yang, Luming, et al.
Published: (2025)
by: Yang, Luming, et al.
Published: (2025)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
by: Shen, Boyang, et al.
Published: (2026)
by: Shen, Boyang, et al.
Published: (2026)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
by: Du, Fan, et al.
Published: (2026)
by: Du, Fan, et al.
Published: (2026)
Multi-scale Temporal Fusion Transformer for Incomplete Vehicle Trajectory Prediction
by: Liu, Zhanwen, et al.
Published: (2024)
by: Liu, Zhanwen, et al.
Published: (2024)
Action-Based ADHD Diagnosis in Video
by: Li, Yichun, et al.
Published: (2024)
by: Li, Yichun, et al.
Published: (2024)
SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation
by: Liu, Qi, et al.
Published: (2023)
by: Liu, Qi, et al.
Published: (2023)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
ReL-SAR: Representation Learning for Skeleton Action Recognition with Convolutional Transformers and BYOL
by: Naimi, Safwen, et al.
Published: (2024)
by: Naimi, Safwen, et al.
Published: (2024)
Interpretable Action Recognition on Hard to Classify Actions
by: Anichenko, Anastasia, et al.
Published: (2024)
by: Anichenko, Anastasia, et al.
Published: (2024)
One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
by: Yang, Siyuan, et al.
Published: (2023)
by: Yang, Siyuan, et al.
Published: (2023)
Object-Centric Latent Action Learning
by: Klepach, Albina, et al.
Published: (2025)
by: Klepach, Albina, et al.
Published: (2025)
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
by: Zhou, Yuhao, et al.
Published: (2026)
by: Zhou, Yuhao, et al.
Published: (2026)
MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction
by: Feng, Yan, et al.
Published: (2024)
by: Feng, Yan, et al.
Published: (2024)
The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024
by: Han, Yinan, et al.
Published: (2024)
by: Han, Yinan, et al.
Published: (2024)
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
by: Zhuang, Qiyuan, et al.
Published: (2026)
by: Zhuang, Qiyuan, et al.
Published: (2026)
Multi-scale Spatio-temporal Transformer-based Imbalanced Longitudinal Learning for Glaucoma Forecasting from Irregular Time Series Images
by: Yang, Xikai, et al.
Published: (2024)
by: Yang, Xikai, et al.
Published: (2024)
Action-Agnostic Point-Level Supervision for Temporal Action Detection
by: Yoshida, Shuhei M., et al.
Published: (2024)
by: Yoshida, Shuhei M., et al.
Published: (2024)
CADIC: Continual Anomaly Detection Based on Incremental Coreset
by: Yang, Gen, et al.
Published: (2025)
by: Yang, Gen, et al.
Published: (2025)
MoST: Motion Style Transformer between Diverse Action Contents
by: Kim, Boeun, et al.
Published: (2024)
by: Kim, Boeun, et al.
Published: (2024)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025)
by: Zahid, Azizul, et al.
Published: (2025)
Mamba Fusion: Learning Actions Through Questioning
by: Dong, Zhikang, et al.
Published: (2024)
by: Dong, Zhikang, et al.
Published: (2024)
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)
by: Garrido, Quentin, et al.
Published: (2026)
Flatten: Video Action Recognition is an Image Classification task
by: Chen, Junlin, et al.
Published: (2024)
by: Chen, Junlin, et al.
Published: (2024)
Learning Vision-Language-Action World Models for Autonomous Driving
by: Wang, Guoqing, et al.
Published: (2026)
by: Wang, Guoqing, et al.
Published: (2026)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
by: Huang, Wei-Jhe, et al.
Published: (2024)
by: Huang, Wei-Jhe, et al.
Published: (2024)
ContextDet: Temporal Action Detection with Adaptive Context Aggregation
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
ActionParty: Multi-Subject Action Binding in Generative Video Games
by: Pondaven, Alexander, et al.
Published: (2026)
by: Pondaven, Alexander, et al.
Published: (2026)
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
by: Ni, Xinzhe, et al.
Published: (2022)
by: Ni, Xinzhe, et al.
Published: (2022)
Variational Contrastive Learning for Skeleton-based Action Recognition
by: Nguyen, Dang Dinh, et al.
Published: (2026)
by: Nguyen, Dang Dinh, et al.
Published: (2026)
Similar Items
-
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024) -
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors
by: Yuan, Kaishen, et al.
Published: (2024) -
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
by: Xie, Liping, et al.
Published: (2025) -
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
by: Liu, Chenyv, et al.
Published: (2026) -
Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic
by: Tang, Jianwei, et al.
Published: (2025)