Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yiming, Zhao, Zhuokai, Chen, Zhaorun, Ding, Zenghui, Yang, Xianjun, Sun, Yining
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912288445825024
author Zhang, Yiming
Zhao, Zhuokai
Chen, Zhaorun
Ding, Zenghui
Yang, Xianjun
Sun, Yining
author_facet Zhang, Yiming
Zhao, Zhuokai
Chen, Zhaorun
Ding, Zenghui
Yang, Xianjun
Sun, Yining
contents Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal details, which incurs significant data and computation costs. In contrast, training-free approaches, though efficient, often lack robustness in preserving context-rich features across complex video content. To this end, we propose DYTO, a novel dynamic token merging framework for zero-shot video understanding that adaptively optimizes token efficiency while preserving crucial scene details. DYTO integrates a hierarchical frame selection and a bipartite token merging strategy to dynamically cluster key frames and selectively compress token sequences, striking a balance between computational efficiency with semantic richness. Extensive experiments across multiple benchmarks demonstrate the effectiveness of DYTO, achieving superior performance compared to both fine-tuned and training-free methods and setting a new state-of-the-art for zero-shot video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14401
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Zhang, Yiming
Zhao, Zhuokai
Chen, Zhaorun
Ding, Zenghui
Yang, Xianjun
Sun, Yining
Computer Vision and Pattern Recognition
Machine Learning
Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal details, which incurs significant data and computation costs. In contrast, training-free approaches, though efficient, often lack robustness in preserving context-rich features across complex video content. To this end, we propose DYTO, a novel dynamic token merging framework for zero-shot video understanding that adaptively optimizes token efficiency while preserving crucial scene details. DYTO integrates a hierarchical frame selection and a bipartite token merging strategy to dynamically cluster key frames and selectively compress token sequences, striking a balance between computational efficiency with semantic richness. Extensive experiments across multiple benchmarks demonstrate the effectiveness of DYTO, achieving superior performance compared to both fine-tuned and training-free methods and setting a new state-of-the-art for zero-shot video understanding.
title Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.14401