A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhuang, Peiqin, Bai, Lei, Wu, Yichao, Liang, Ding, Zhou, Luping, Wang, Yali, Ouyang, Wanli |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
by: Su, Rui, et al.
Published: (2025)
by: Su, Rui, et al.
Published: (2025)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)
by: Yue, Xiaoyu, et al.
Published: (2025)
Diffusion Models Need Visual Priors for Image Generation
by: Yue, Xiaoyu, et al.
Published: (2024)
by: Yue, Xiaoyu, et al.
Published: (2024)
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
LORTSAR: Low-Rank Transformer for Skeleton-based Action Recognition
by: Oraki, Soroush, et al.
Published: (2024)
by: Oraki, Soroush, et al.
Published: (2024)
Low-Resolution Action Recognition for Tiny Actions Challenge
by: Chen, Boyu, et al.
Published: (2022)
by: Chen, Boyu, et al.
Published: (2022)
FiT: Flexible Vision Transformer for Diffusion Model
by: Lu, Zeyu, et al.
Published: (2024)
by: Lu, Zeyu, et al.
Published: (2024)
Action Recognition with Multi-stream Motion Modeling and Mutual Information Maximization
by: Yang, Yuheng, et al.
Published: (2023)
by: Yang, Yuheng, et al.
Published: (2023)
TCFormer: Visual Recognition via Token Clustering Transformer
by: Zeng, Wang, et al.
Published: (2024)
by: Zeng, Wang, et al.
Published: (2024)
SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action Recognition
by: Huang, Wenbo, et al.
Published: (2024)
by: Huang, Wenbo, et al.
Published: (2024)
Evolving Skeletons: Motion Dynamics in Action Recognition
by: Qiu, Jushang, et al.
Published: (2025)
by: Qiu, Jushang, et al.
Published: (2025)
Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
Fast Information Streaming Handler (FisH): A Unified Seismic Neural Network for Single Station Real-Time Earthquake Early Warning
by: Zhang, Tianning, et al.
Published: (2024)
by: Zhang, Tianning, et al.
Published: (2024)
X-Edit: Exact, Explicit, and Explainable Null-Space Editing for Medical Vision Transformers
by: Liu, Yuanye, et al.
Published: (2026)
by: Liu, Yuanye, et al.
Published: (2026)
LOCR: Location-Guided Transformer for Optical Character Recognition
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
by: Wu, Wenhao, et al.
Published: (2023)
by: Wu, Wenhao, et al.
Published: (2023)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
by: Zhang, Yaqi, et al.
Published: (2023)
by: Zhang, Yaqi, et al.
Published: (2023)
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
Native-Resolution Image Synthesis
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
Taylor Videos for Action Recognition
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026)
by: Kuang, Jidong, et al.
Published: (2026)
Human-Centric Transformer for Domain Adaptive Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
Analysis and Evaluation of Kinect-based Action Recognition Algorithms
by: Wang, Lei
Published: (2021)
by: Wang, Lei
Published: (2021)
Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
by: Tu, Chongjun, et al.
Published: (2025)
by: Tu, Chongjun, et al.
Published: (2025)
Explicit Interaction for Fusion-Based Place Recognition
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
ReconMOST: Multi-Layer Sea Temperature Reconstruction with Observations-Guided Diffusion
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
VegeDiff: Latent Diffusion Model for Geospatial Vegetation Forecasting
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
RS-Mamba for Large Remote Sensing Image Dense Prediction
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
ALow-Cost Real-Time Framework for Industrial Action Recognition Using Foundation Models
by: Wang, Zhicheng, et al.
Published: (2024)
by: Wang, Zhicheng, et al.
Published: (2024)
Disentangling Static and Dynamic Information for Reducing Static Bias in Action Recognition
by: Kobayashi, Masato, et al.
Published: (2025)
by: Kobayashi, Masato, et al.
Published: (2025)
Transforming Weather Data from Pixel to Latent Space
by: Zhao, Sijie, et al.
Published: (2025)
by: Zhao, Sijie, et al.
Published: (2025)
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
Prompt-guided Disentangled Representation for Action Recognition
by: Wu, Tianci, et al.
Published: (2025)
by: Wu, Tianci, et al.
Published: (2025)
PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis
by: Wang, Zicheng, et al.
Published: (2024)
by: Wang, Zicheng, et al.
Published: (2024)
PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition
by: Wang, Jie, et al.
Published: (2025)
by: Wang, Jie, et al.
Published: (2025)
Kinetic Mining in Context: Few-Shot Action Synthesis via Text-to-Motion Distillation
by: Cazzola, Luca, et al.
Published: (2025)
by: Cazzola, Luca, et al.
Published: (2025)
STAR: A Benchmark for Astronomical Star Fields Super-Resolution
by: Wu, Kuo-Cheng, et al.
Published: (2025)
by: Wu, Kuo-Cheng, et al.
Published: (2025)
SVFormer: A Direct Training Spiking Transformer for Efficient Video Action Recognition
by: Yu, Liutao, et al.
Published: (2024)
by: Yu, Liutao, et al.
Published: (2024)
Similar Items
-
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
by: Su, Rui, et al.
Published: (2025) -
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019) -
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025) -
Diffusion Models Need Visual Priors for Image Generation
by: Yue, Xiaoyu, et al.
Published: (2024) -
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)