OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kang, Minseok, Lee, Minhyeok, Lee, Jungho, Kim, Minjung, Kim, Donghyeong, Lee, Dayeon, Choi, Heeseung, Kim, Ig-jae, Lee, Sangyoun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916004319199232
author Kang, Minseok
Lee, Minhyeok
Lee, Jungho
Kim, Minjung
Kim, Donghyeong
Lee, Dayeon
Choi, Heeseung
Kim, Ig-jae
Lee, Sangyoun
author_facet Kang, Minseok
Lee, Minhyeok
Lee, Jungho
Kim, Minjung
Kim, Donghyeong
Lee, Dayeon
Choi, Heeseung
Kim, Ig-jae
Lee, Sangyoun
contents As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across frames. Training-free token compression has emerged as a practical solution to this bottleneck. However, existing temporal compression methods rely primarily on cross-frame token similarity or segmentation heuristics, overlooking each token's semantic role within its frame and failing to adapt compression strength to the compressibility of each frame pair. In this work, we propose OTT-Vid, a transport-derived allocation framework for temporal token compression. Our approach consists of two stages: spatial pruning identifies representative content within each frame, and optimal transport (OT) is then solved between neighboring frames to estimate temporal compressibility. We formulate this OT with non-uniform token mass, which protects semantically important tokens from aggressive compression, and a locality-aware cost that captures both feature and spatial disparities. The resulting transport plan jointly balances token importance and matching cost, while its total cost defines the transport difficulty of each frame pair, which we use to allocate compression budgets dynamically. Experiments on six benchmarks spanning video question answering and temporal grounding show that OTT-Vid preserves 95.8% of VQA and 73.9% of VTG performance while retaining only 10% of tokens, consistently outperforming existing state-of-the-art training-free compression methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11803
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
Kang, Minseok
Lee, Minhyeok
Lee, Jungho
Kim, Minjung
Kim, Donghyeong
Lee, Dayeon
Choi, Heeseung
Kim, Ig-jae
Lee, Sangyoun
Computer Vision and Pattern Recognition
Artificial Intelligence
As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across frames. Training-free token compression has emerged as a practical solution to this bottleneck. However, existing temporal compression methods rely primarily on cross-frame token similarity or segmentation heuristics, overlooking each token's semantic role within its frame and failing to adapt compression strength to the compressibility of each frame pair. In this work, we propose OTT-Vid, a transport-derived allocation framework for temporal token compression. Our approach consists of two stages: spatial pruning identifies representative content within each frame, and optimal transport (OT) is then solved between neighboring frames to estimate temporal compressibility. We formulate this OT with non-uniform token mass, which protects semantically important tokens from aggressive compression, and a locality-aware cost that captures both feature and spatial disparities. The resulting transport plan jointly balances token importance and matching cost, while its total cost defines the transport difficulty of each frame pair, which we use to allocate compression budgets dynamically. Experiments on six benchmarks spanning video question answering and temporal grounding show that OTT-Vid preserves 95.8% of VQA and 73.9% of VTG performance while retaining only 10% of tokens, consistently outperforming existing state-of-the-art training-free compression methods.
title OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.11803