Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Pulkit, Padmanabhan, Namitha, Luo, Luke, Rambhatla, Sai Saketh, Shrivastava, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)
by: Kumar, Pulkit, et al.
Published: (2025)
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026)
by: Padmanabhan, Namitha, et al.
Published: (2026)
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024)
by: Padmanabhan, Namitha, et al.
Published: (2024)
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
SelfEval: Leveraging the discriminative nature of generative models for evaluation
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
How to Design and Train Your Implicit Neural Representation for Video Compression
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026)
by: Agarwal, Vatsal, et al.
Published: (2026)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
A Comprehensive Review of Few-shot Action Recognition
by: Wanyan, Yuyang, et al.
Published: (2024)
by: Wanyan, Yuyang, et al.
Published: (2024)
Hierarchical Compositional Representations for Few-shot Action Recognition
by: Li, Changzhen, et al.
Published: (2022)
by: Li, Changzhen, et al.
Published: (2022)
Do text-free diffusion models learn discriminative visual representations?
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
CLIP-guided Prototype Modulating for Few-shot Action Recognition
by: Wang, Xiang, et al.
Published: (2023)
by: Wang, Xiang, et al.
Published: (2023)
Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
by: Qian, Zefeng, et al.
Published: (2025)
by: Qian, Zefeng, et al.
Published: (2025)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Active Multimodal Distillation for Few-shot Action Recognition
by: Feng, Weijia, et al.
Published: (2025)
by: Feng, Weijia, et al.
Published: (2025)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Unified Framework for Open-World Compositional Zero-shot Learning
by: Jayasekara, Hirunima, et al.
Published: (2024)
by: Jayasekara, Hirunima, et al.
Published: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition
by: Qian, Zefeng, et al.
Published: (2025)
by: Qian, Zefeng, et al.
Published: (2025)
Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action Recognition
by: Cao, Congqi, et al.
Published: (2024)
by: Cao, Congqi, et al.
Published: (2024)
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
Adaptive Prototype Model for Attribute-based Multi-label Few-shot Action Recognition
by: Xiao, Juefeng, et al.
Published: (2025)
by: Xiao, Juefeng, et al.
Published: (2025)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition
by: Cao, Congqi, et al.
Published: (2025)
by: Cao, Congqi, et al.
Published: (2025)
Beyond Class Tokens: LLM-guided Dominant Property Mining for Few-shot Classification
by: Zhuo, Wei, et al.
Published: (2025)
by: Zhuo, Wei, et al.
Published: (2025)
Scale Space Diffusion
by: Mukhopadhyay, Soumik, et al.
Published: (2026)
by: Mukhopadhyay, Soumik, et al.
Published: (2026)
DMSD-CDFSAR: Distillation from Mixed-Source Domain for Cross-Domain Few-shot Action Recognition
by: Guo, Fei, et al.
Published: (2024)
by: Guo, Fei, et al.
Published: (2024)
D$^2$ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition
by: Pei, Wenjie, et al.
Published: (2023)
by: Pei, Wenjie, et al.
Published: (2023)
Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning
by: Yi, Shuai, et al.
Published: (2025)
by: Yi, Shuai, et al.
Published: (2025)
MotiF: Making Text Count in Image Animation with Motion Focal Loss
by: Wang, Shijie, et al.
Published: (2024)
by: Wang, Shijie, et al.
Published: (2024)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Appearance-free Action Recognition: Zero-shot Generalization in Humans and a Two-Pathway Model
by: Kumar, Prerana, et al.
Published: (2026)
by: Kumar, Prerana, et al.
Published: (2026)
Multi-view Distillation based on Multi-modal Fusion for Few-shot Action Recognition(CLIP-$\mathrm{M^2}$DF)
by: Guo, Fei, et al.
Published: (2024)
by: Guo, Fei, et al.
Published: (2024)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024)
by: Wang, Hanyu, et al.
Published: (2024)
V-VIPE: Variational View Invariant Pose Embedding
by: Levy, Mara, et al.
Published: (2024)
by: Levy, Mara, et al.
Published: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Zero-shot Compositional Action Recognition with Neural Logic Constraints
by: Ye, Gefan, et al.
Published: (2025)
by: Ye, Gefan, et al.
Published: (2025)
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025)
by: Aggarwal, Anirud, et al.
Published: (2025)
Mitigating Hallucinations in Diffusion Models through Adaptive Attention Modulation
by: Oorloff, Trevine, et al.
Published: (2025)
by: Oorloff, Trevine, et al.
Published: (2025)
Robust Saliency-Aware Distillation for Few-shot Fine-grained Visual Recognition
by: Liu, Haiqi, et al.
Published: (2023)
by: Liu, Haiqi, et al.
Published: (2023)
Similar Items
-
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025) -
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026) -
Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing their Contributions
by: Padmanabhan, Namitha, et al.
Published: (2024) -
UVIS: Unsupervised Video Instance Segmentation
by: Huang, Shuaiyi, et al.
Published: (2024) -
SelfEval: Leveraging the discriminative nature of generative models for evaluation
by: Rambhatla, Sai Saketh, et al.
Published: (2023)