Salvato in:
| Autori principali: | Moodley, Perusha, Kaushik, Pramod, Thambi, Dhillu, Trovinger, Mark, Paruchuri, Praveen, Hong, Xia, Rosman, Benjamin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.01310 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MV-GMN: State Space Model for Multi-View Action Recognition
di: Lin, Yuhui, et al.
Pubblicazione: (2025)
di: Lin, Yuhui, et al.
Pubblicazione: (2025)
MALT: Multi-scale Action Learning Transformer for Online Action Detection
di: Yang, Zhipeng, et al.
Pubblicazione: (2024)
di: Yang, Zhipeng, et al.
Pubblicazione: (2024)
Action Selection Learning for Multi-label Multi-view Action Recognition
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2024)
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
di: Chandra, Soumyadeep, et al.
Pubblicazione: (2024)
di: Chandra, Soumyadeep, et al.
Pubblicazione: (2024)
S3T-Former: A Purely Spike-Driven State-Space Topology Transformer for Skeleton Action Recognition
di: Zheng, Naichuan, et al.
Pubblicazione: (2026)
di: Zheng, Naichuan, et al.
Pubblicazione: (2026)
Boundary Discretization and Reliable Classification Network for Temporal Action Detection
di: Fang, Zhenying, et al.
Pubblicazione: (2023)
di: Fang, Zhenying, et al.
Pubblicazione: (2023)
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2025)
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2025)
A Real-Time Human Action Recognition Model for Assisted Living
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
di: Kiray, Mert, et al.
Pubblicazione: (2025)
di: Kiray, Mert, et al.
Pubblicazione: (2025)
A Universal Action Space for General Behavior Analysis
di: Chang, Hung-Shuo, et al.
Pubblicazione: (2026)
di: Chang, Hung-Shuo, et al.
Pubblicazione: (2026)
Multi-Granularity Hand Action Detection
di: Zhe, Ting, et al.
Pubblicazione: (2023)
di: Zhe, Ting, et al.
Pubblicazione: (2023)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
di: Yamane, Taiga, et al.
Pubblicazione: (2025)
di: Yamane, Taiga, et al.
Pubblicazione: (2025)
Hierarchical Multi-Stage Transformer Architecture for Context-Aware Temporal Action Localization
di: Ullah, Hayat, et al.
Pubblicazione: (2025)
di: Ullah, Hayat, et al.
Pubblicazione: (2025)
Learning Action Hierarchies via Hybrid Geometric Diffusion
di: Kaushik, Arjun Ramesh, et al.
Pubblicazione: (2026)
di: Kaushik, Arjun Ramesh, et al.
Pubblicazione: (2026)
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025)
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025)
Multi-level and Multi-modal Action Anticipation
di: Kim, Seulgi, et al.
Pubblicazione: (2025)
di: Kim, Seulgi, et al.
Pubblicazione: (2025)
Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
di: Shuvo, Rezowan, et al.
Pubblicazione: (2025)
di: Shuvo, Rezowan, et al.
Pubblicazione: (2025)
MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
di: Fang, Zhenying, et al.
Pubblicazione: (2025)
di: Fang, Zhenying, et al.
Pubblicazione: (2025)
MultiModal Action Conditioned Video Generation
di: Li, Yichen, et al.
Pubblicazione: (2025)
di: Li, Yichen, et al.
Pubblicazione: (2025)
Towards Scalable Modeling of Compressed Videos for Efficient Action Recognition
di: Biswas, Shristi Das, et al.
Pubblicazione: (2025)
di: Biswas, Shristi Das, et al.
Pubblicazione: (2025)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
di: Liang, Zhixuan, et al.
Pubblicazione: (2025)
di: Liang, Zhixuan, et al.
Pubblicazione: (2025)
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2025)
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
Friends Across Time: Multi-Scale Action Segmentation Transformer for Surgical Phase Recognition
di: Zhang, Bokai, et al.
Pubblicazione: (2024)
di: Zhang, Bokai, et al.
Pubblicazione: (2024)
Multi-Stage Contrastive Regression for Action Quality Assessment
di: An, Qi, et al.
Pubblicazione: (2024)
di: An, Qi, et al.
Pubblicazione: (2024)
Dual DETRs for Multi-Label Temporal Action Detection
di: Zhu, Yuhan, et al.
Pubblicazione: (2024)
di: Zhu, Yuhan, et al.
Pubblicazione: (2024)
MMAD: Multi-label Micro-Action Detection in Videos
di: Li, Kun, et al.
Pubblicazione: (2024)
di: Li, Kun, et al.
Pubblicazione: (2024)
Multi-task Learning For Joint Action and Gesture Recognition
di: Spathis, Konstantinos, et al.
Pubblicazione: (2025)
di: Spathis, Konstantinos, et al.
Pubblicazione: (2025)
Advancing Compressed Video Action Recognition through Progressive Knowledge Distillation
di: Soufleri, Efstathia, et al.
Pubblicazione: (2024)
di: Soufleri, Efstathia, et al.
Pubblicazione: (2024)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
di: Souček, Tomáš, et al.
Pubblicazione: (2023)
di: Souček, Tomáš, et al.
Pubblicazione: (2023)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2024)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation
di: Liu, Qi, et al.
Pubblicazione: (2023)
di: Liu, Qi, et al.
Pubblicazione: (2023)
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
di: Wang, Zirui, et al.
Pubblicazione: (2024)
di: Wang, Zirui, et al.
Pubblicazione: (2024)
Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent Space
di: Nakagawa, Ren, et al.
Pubblicazione: (2025)
di: Nakagawa, Ren, et al.
Pubblicazione: (2025)
An Effective-Efficient Approach for Dense Multi-Label Action Detection
di: Sardari, Faegheh, et al.
Pubblicazione: (2024)
di: Sardari, Faegheh, et al.
Pubblicazione: (2024)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
di: Peng, Liyang, et al.
Pubblicazione: (2025)
di: Peng, Liyang, et al.
Pubblicazione: (2025)
MAMMA: Markerless & Automatic Multi-Person Motion Action Capture
di: Cuevas-Velasquez, Hanz, et al.
Pubblicazione: (2025)
di: Cuevas-Velasquez, Hanz, et al.
Pubblicazione: (2025)
ActionParty: Multi-Subject Action Binding in Generative Video Games
di: Pondaven, Alexander, et al.
Pubblicazione: (2026)
di: Pondaven, Alexander, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MV-GMN: State Space Model for Multi-View Action Recognition
di: Lin, Yuhui, et al.
Pubblicazione: (2025) -
MALT: Multi-scale Action Learning Transformer for Online Action Detection
di: Yang, Zhipeng, et al.
Pubblicazione: (2024) -
Action Selection Learning for Multi-label Multi-view Action Recognition
di: Nguyen, Trung Thanh, et al.
Pubblicazione: (2024) -
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
di: Chandra, Soumyadeep, et al.
Pubblicazione: (2024) -
S3T-Former: A Purely Spike-Driven State-Space Topology Transformer for Skeleton Action Recognition
di: Zheng, Naichuan, et al.
Pubblicazione: (2026)