One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Chenhao, Zhang, Jieyu, Salehi, Mohammadreza, Gao, Ziqi, Iyengar, Vishnu, Kobori, Norimasa, Kong, Quan, Krishna, Ranjay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911665405034496
author Zheng, Chenhao
Zhang, Jieyu
Salehi, Mohammadreza
Gao, Ziqi
Iyengar, Vishnu
Kobori, Norimasa
Kong, Quan
Krishna, Ranjay
author_facet Zheng, Chenhao
Zhang, Jieyu
Salehi, Mohammadreza
Gao, Ziqi
Iyengar, Vishnu
Kobori, Norimasa
Kong, Quan
Krishna, Ranjay
contents Effective video tokenization is critical for scaling transformer models for long videos. Current approaches tokenize videos using space-time patches, leading to excessive tokens and computational inefficiencies. The best token reduction strategies degrade performance and barely reduce the number of tokens when the camera moves. We introduce grounded video tokenization, a paradigm that organizes tokens based on panoptic sub-object trajectories rather than fixed patches. Our method aligns with fundamental perceptual principles, ensuring that tokenization reflects scene complexity rather than video duration. We propose TrajViT, a video encoder that extracts object trajectories and converts them into semantically meaningful tokens, significantly reducing redundancy while maintaining temporal coherence. Trained with contrastive learning, TrajViT significantly outperforms space-time ViT (ViT3D) across multiple video understanding benchmarks, e.g., TrajViT outperforms ViT3D by a large margin of 6% top-5 recall in average at video-text retrieval task with 10x token deduction. We also show TrajViT as a stronger model than ViT3D for being the video encoder for modern VideoLLM, obtaining an average of 5.2% performance improvement across 6 VideoQA benchmarks while having 4x faster training time and 18x less inference FLOPs. TrajViT is the first efficient encoder to consistently outperform ViT3D across diverse video analysis tasks, making it a robust and scalable solution.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23617
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
Zheng, Chenhao
Zhang, Jieyu
Salehi, Mohammadreza
Gao, Ziqi
Iyengar, Vishnu
Kobori, Norimasa
Kong, Quan
Krishna, Ranjay
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Effective video tokenization is critical for scaling transformer models for long videos. Current approaches tokenize videos using space-time patches, leading to excessive tokens and computational inefficiencies. The best token reduction strategies degrade performance and barely reduce the number of tokens when the camera moves. We introduce grounded video tokenization, a paradigm that organizes tokens based on panoptic sub-object trajectories rather than fixed patches. Our method aligns with fundamental perceptual principles, ensuring that tokenization reflects scene complexity rather than video duration. We propose TrajViT, a video encoder that extracts object trajectories and converts them into semantically meaningful tokens, significantly reducing redundancy while maintaining temporal coherence. Trained with contrastive learning, TrajViT significantly outperforms space-time ViT (ViT3D) across multiple video understanding benchmarks, e.g., TrajViT outperforms ViT3D by a large margin of 6% top-5 recall in average at video-text retrieval task with 10x token deduction. We also show TrajViT as a stronger model than ViT3D for being the video encoder for modern VideoLLM, obtaining an average of 5.2% performance improvement across 6 VideoQA benchmarks while having 4x faster training time and 18x less inference FLOPs. TrajViT is the first efficient encoder to consistently outperform ViT3D across diverse video analysis tasks, making it a robust and scalable solution.
title One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2505.23617