Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yiming, Zhao, Zhuokai, Chen, Zhaorun, Ding, Zenghui, Yang, Xianjun, Sun, Yining |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RankCLIP: Ranking-Consistent Language-Image Pretraining
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
by: Fang, Yixiong, et al.
Published: (2024)
by: Fang, Yixiong, et al.
Published: (2024)
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
HyperTokens: Controlling Token Dynamics for Continual Video-Language Understanding
by: Nguyen, Toan, et al.
Published: (2026)
by: Nguyen, Toan, et al.
Published: (2026)
Video Token Merging for Long-form Video Understanding
by: Lee, Seon-Ho, et al.
Published: (2024)
by: Lee, Seon-Ho, et al.
Published: (2024)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
by: Beliaev, Mark, et al.
Published: (2025)
by: Beliaev, Mark, et al.
Published: (2025)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
by: Jiang, Xixi, et al.
Published: (2025)
by: Jiang, Xixi, et al.
Published: (2025)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
by: Cao, Cong, et al.
Published: (2024)
by: Cao, Cong, et al.
Published: (2024)
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
by: Liu, Kaiting, et al.
Published: (2026)
by: Liu, Kaiting, et al.
Published: (2026)
SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Efficient Visual Transformer by Learnable Token Merging
by: Wang, Yancheng, et al.
Published: (2024)
by: Wang, Yancheng, et al.
Published: (2024)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
by: Shen, Leqi, et al.
Published: (2024)
by: Shen, Leqi, et al.
Published: (2024)
Language-Driven Anchors for Zero-Shot Adversarial Robustness
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
Zero-Shot Video Translation via Token Warping
by: Zhu, Haiming, et al.
Published: (2024)
by: Zhu, Haiming, et al.
Published: (2024)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
by: Jiang, Shuyang, et al.
Published: (2026)
by: Jiang, Shuyang, et al.
Published: (2026)
Navigate Beyond Shortcuts: Debiased Learning through the Lens of Neural Collapse
by: Wang, Yining, et al.
Published: (2024)
by: Wang, Yining, et al.
Published: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
by: Gong, Wenyi, et al.
Published: (2025)
by: Gong, Wenyi, et al.
Published: (2025)
Zero-Shot Long-Form Video Understanding through Screenplay
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025)
by: Zhang, Siyang, et al.
Published: (2025)
Are Image-to-Video Models Good Zero-Shot Image Editors?
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)
by: Zhang, Luyuan, et al.
Published: (2026)
Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers
by: Wang, Hongjie, et al.
Published: (2023)
by: Wang, Hongjie, et al.
Published: (2023)
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
by: Fan, Ziyang, et al.
Published: (2026)
by: Fan, Ziyang, et al.
Published: (2026)
Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues
by: Zambare, Pallavi, et al.
Published: (2025)
by: Zambare, Pallavi, et al.
Published: (2025)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
Video, How Do Your Tokens Merge?
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
by: Chang, Shuochen, et al.
Published: (2025)
by: Chang, Shuochen, et al.
Published: (2025)
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner
by: Zhou, Yufan, et al.
Published: (2024)
by: Zhou, Yufan, et al.
Published: (2024)
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
by: Tao, Keda, et al.
Published: (2024)
by: Tao, Keda, et al.
Published: (2024)
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
by: Zhang, Song, et al.
Published: (2026)
by: Zhang, Song, et al.
Published: (2026)
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
by: Zhao, Zelin, et al.
Published: (2025)
by: Zhao, Zelin, et al.
Published: (2025)
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
by: Abdal, Rameen, et al.
Published: (2025)
by: Abdal, Rameen, et al.
Published: (2025)
CodeMerge: Codebook-Guided Model Merging for Robust Test-Time Adaptation in Autonomous Driving
by: Yang, Huitong, et al.
Published: (2025)
by: Yang, Huitong, et al.
Published: (2025)
Domain-Aware Continual Zero-Shot Learning
by: Yi, Kai, et al.
Published: (2021)
by: Yi, Kai, et al.
Published: (2021)
Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators
by: Zhang, Wentao, et al.
Published: (2024)
by: Zhang, Wentao, et al.
Published: (2024)
ShoulderShot: Generating Over-the-Shoulder Dialogue Videos
by: Zhang, Yuang, et al.
Published: (2025)
by: Zhang, Yuang, et al.
Published: (2025)
Similar Items
-
RankCLIP: Ranking-Consistent Language-Image Pretraining
by: Zhang, Yiming, et al.
Published: (2024) -
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
by: Zhang, Yiming, et al.
Published: (2025) -
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
by: Fang, Yixiong, et al.
Published: (2024) -
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
by: Chen, Zhaorun, et al.
Published: (2024) -
HyperTokens: Controlling Token Dynamics for Continual Video-Language Understanding
by: Nguyen, Toan, et al.
Published: (2026)