InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917356000772096 |
|---|---|
| author | Ye, Haotian He, Qiyuan Han, Jiaqi Li, Puheng Fan, Jiaojiao Hao, Zekun Reda, Fitsum Balaji, Yogesh Chen, Huayu Liu, Sheng Yao, Angela Zou, James Ermon, Stefano Wang, Haoxiang Liu, Ming-Yu |
| author_facet | Ye, Haotian He, Qiyuan Han, Jiaqi Li, Puheng Fan, Jiaojiao Hao, Zekun Reda, Fitsum Balaji, Yogesh Chen, Huayu Liu, Sheng Yao, Angela Zou, James Ermon, Stefano Wang, Haoxiang Liu, Ming-Yu |
| contents | Accurate and efficient discrete video tokenization is essential for long video sequences processing. Yet, the inherent complexity and variable information density of videos present a significant bottleneck for current tokenizers, which rigidly compress all content at a fixed rate, leading to redundancy or information loss. Drawing inspiration from Shannon's information theory, this paper introduces InfoTok, a principled framework for adaptive video tokenization. We rigorously prove that existing data-agnostic training methods are suboptimal in representation length, and present a novel evidence lower bound (ELBO)-based algorithm that approaches theoretical optimality. Leveraging this framework, we develop a transformer-based adaptive compressor that enables adaptive tokenization. Empirical results demonstrate state-of-the-art compression performance, saving 20% tokens without influence on performance, and achieving 2.3x compression rates while still outperforming prior heuristic adaptive approaches. By allocating tokens according to informational richness, InfoTok enables a more compressed yet accurate tokenization for video representation, offering valuable insights for future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_16975 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression Ye, Haotian He, Qiyuan Han, Jiaqi Li, Puheng Fan, Jiaojiao Hao, Zekun Reda, Fitsum Balaji, Yogesh Chen, Huayu Liu, Sheng Yao, Angela Zou, James Ermon, Stefano Wang, Haoxiang Liu, Ming-Yu Computer Vision and Pattern Recognition Artificial Intelligence Accurate and efficient discrete video tokenization is essential for long video sequences processing. Yet, the inherent complexity and variable information density of videos present a significant bottleneck for current tokenizers, which rigidly compress all content at a fixed rate, leading to redundancy or information loss. Drawing inspiration from Shannon's information theory, this paper introduces InfoTok, a principled framework for adaptive video tokenization. We rigorously prove that existing data-agnostic training methods are suboptimal in representation length, and present a novel evidence lower bound (ELBO)-based algorithm that approaches theoretical optimality. Leveraging this framework, we develop a transformer-based adaptive compressor that enables adaptive tokenization. Empirical results demonstrate state-of-the-art compression performance, saving 20% tokens without influence on performance, and achieving 2.3x compression rates while still outperforming prior heuristic adaptive approaches. By allocating tokens according to informational richness, InfoTok enables a more compressed yet accurate tokenization for video representation, offering valuable insights for future research. |
| title | InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2512.16975 |