Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Pulkit, Huang, Shuaiyi, Walmer, Matthew, Rambhatla, Sai Saketh, Shrivastava, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2024)
by: Kumar, Pulkit, et al.
Published: (2024)
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
SelfEval: Leveraging the discriminative nature of generative models for evaluation
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026)
by: Walmer, Matthew, et al.
Published: (2026)
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)
by: Suri, Saksham, et al.
Published: (2024)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026)
by: Agarwal, Vatsal, et al.
Published: (2026)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
Multi-entity Video Transformers for Fine-Grained Video Representation Learning
by: Walmer, Matthew, et al.
Published: (2023)
by: Walmer, Matthew, et al.
Published: (2023)
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024)
by: Padmanabhan, Namitha, et al.
Published: (2024)
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations
by: Huang, Shuaiyi, et al.
Published: (2025)
by: Huang, Shuaiyi, et al.
Published: (2025)
TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition
by: Wang, Yilong, et al.
Published: (2024)
by: Wang, Yilong, et al.
Published: (2024)
MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2024)
by: Qu, Hongyu, et al.
Published: (2024)
Dual-View Data Hallucination with Semantic Relation Guidance for Few-Shot Image Recognition
by: Wu, Hefeng, et al.
Published: (2024)
by: Wu, Hefeng, et al.
Published: (2024)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
by: Li, Bozheng, et al.
Published: (2024)
by: Li, Bozheng, et al.
Published: (2024)
STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
by: Liu, Hongli, et al.
Published: (2026)
by: Liu, Hongli, et al.
Published: (2026)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
by: Qian, Zefeng, et al.
Published: (2025)
by: Qian, Zefeng, et al.
Published: (2025)
SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action Recognition
by: Huang, Wenbo, et al.
Published: (2024)
by: Huang, Wenbo, et al.
Published: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence
by: Huang, Wenbo, et al.
Published: (2024)
by: Huang, Wenbo, et al.
Published: (2024)
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2026)
by: Qu, Hongyu, et al.
Published: (2026)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition
by: Xing, Jiazheng, et al.
Published: (2023)
by: Xing, Jiazheng, et al.
Published: (2023)
Task-Specific Distance Correlation Matching for Few-Shot Action Recognition
by: Long, Fei, et al.
Published: (2025)
by: Long, Fei, et al.
Published: (2025)
Novel Semantic Prompting for Zero-Shot Action Recognition
by: Iqbal, Salman, et al.
Published: (2026)
by: Iqbal, Salman, et al.
Published: (2026)
What is Point Supervision Worth in Video Instance Segmentation?
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
ARDuP: Active Region Video Diffusion for Universal Policies
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025)
by: Aggarwal, Anirud, et al.
Published: (2025)
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026)
by: Padmanabhan, Namitha, et al.
Published: (2026)
Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action Recognition
by: Li, Yuke, et al.
Published: (2024)
by: Li, Yuke, et al.
Published: (2024)
Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
by: Huang, Wenbo, et al.
Published: (2025)
by: Huang, Wenbo, et al.
Published: (2025)
Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
by: Maiya, Shishira R, et al.
Published: (2024)
by: Maiya, Shishira R, et al.
Published: (2024)
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
by: Ni, Xinzhe, et al.
Published: (2022)
by: Ni, Xinzhe, et al.
Published: (2022)
SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning
by: Lai, Jinxiang, et al.
Published: (2023)
by: Lai, Jinxiang, et al.
Published: (2023)
LAD-Drive: Bridging Language and Trajectory with Action-Aware Diffusion Transformers
by: Schmidt, Fabian, et al.
Published: (2026)
by: Schmidt, Fabian, et al.
Published: (2026)
Video-to-Task Learning via Motion-Guided Attention for Few-Shot Action Recognition
by: Guo, Hanyu, et al.
Published: (2024)
by: Guo, Hanyu, et al.
Published: (2024)
Understanding the Cross-Domain Capabilities of Video-Based Few-Shot Action Recognition Models
by: Markham, Georgia, et al.
Published: (2024)
by: Markham, Georgia, et al.
Published: (2024)
Learning by Neighbor-Aware Semantics, Deciding by Open-form Flows: Towards Robust Zero-Shot Skeleton Action Recognition
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
Context-Aware Pesticide Recommendation via Few-Shot Pest Recognition for Precision Agriculture
by: Ghosh, Anirudha, et al.
Published: (2026)
by: Ghosh, Anirudha, et al.
Published: (2026)
Similar Items
-
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2024) -
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024) -
SelfEval: Leveraging the discriminative nature of generative models for evaluation
by: Rambhatla, Sai Saketh, et al.
Published: (2023) -
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026) -
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)