Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Pulkit, Huang, Shuaiyi, Walmer, Matthew, Rambhatla, Sai Saketh, Shrivastava, Abhinav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908479514476544
author Kumar, Pulkit
Huang, Shuaiyi
Walmer, Matthew
Rambhatla, Sai Saketh
Shrivastava, Abhinav
author_facet Kumar, Pulkit
Huang, Shuaiyi
Walmer, Matthew
Rambhatla, Sai Saketh
Shrivastava, Abhinav
contents Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to track and effectively modeling their motion patterns. We present Trokens, a novel approach that transforms trajectory points into semantic-aware relational tokens for action recognition. First, we introduce a semantic-aware sampling strategy to adaptively distribute tracking points based on object scale and semantic relevance. Second, we develop a motion modeling framework that captures both intra-trajectory dynamics through the Histogram of Oriented Displacements (HoD) and inter-trajectory relationships to model complex action patterns. Our approach effectively combines these trajectory tokens with semantic features to enhance appearance features with motion information, achieving state-of-the-art performance across six diverse few-shot action recognition benchmarks: Something-Something-V2 (both full and small splits), Kinetics, UCF101, HMDB51, and FineGym. For project page see https://trokens-iccv25.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2508_03695
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
Kumar, Pulkit
Huang, Shuaiyi
Walmer, Matthew
Rambhatla, Sai Saketh
Shrivastava, Abhinav
Computer Vision and Pattern Recognition
Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to track and effectively modeling their motion patterns. We present Trokens, a novel approach that transforms trajectory points into semantic-aware relational tokens for action recognition. First, we introduce a semantic-aware sampling strategy to adaptively distribute tracking points based on object scale and semantic relevance. Second, we develop a motion modeling framework that captures both intra-trajectory dynamics through the Histogram of Oriented Displacements (HoD) and inter-trajectory relationships to model complex action patterns. Our approach effectively combines these trajectory tokens with semantic features to enhance appearance features with motion information, achieving state-of-the-art performance across six diverse few-shot action recognition benchmarks: Something-Something-V2 (both full and small splits), Kinetics, UCF101, HMDB51, and FineGym. For project page see https://trokens-iccv25.github.io
title Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.03695