Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hyun, Jeongseok, Hwang, Sukjun, Han, Su Ho, Kim, Taeoh, Lee, Inwoong, Wee, Dongyoon, Lee, Joon-Young, Kim, Seon Joo, Shim, Minho |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025)
by: Han, Su Ho, et al.
Published: (2025)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
Classification Matters: Improving Video Action Detection with Class-Specific Attention
by: Lee, Jinsung, et al.
Published: (2024)
by: Lee, Jinsung, et al.
Published: (2024)
Video-Oasis: Rethinking Evaluation of Video Understanding
by: Lim, Geuntaek, et al.
Published: (2026)
by: Lim, Geuntaek, et al.
Published: (2026)
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images
by: Lee, Jungho, et al.
Published: (2025)
by: Lee, Jungho, et al.
Published: (2025)
CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images
by: Lee, Jungho, et al.
Published: (2024)
by: Lee, Jungho, et al.
Published: (2024)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
by: Lee, Hangyeol, et al.
Published: (2026)
by: Lee, Hangyeol, et al.
Published: (2026)
ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
by: Kang, Hyolim, et al.
Published: (2024)
by: Kang, Hyolim, et al.
Published: (2024)
A Simple Baseline with Single-encoder for Referring Image Segmentation
by: Yu, Seonghoon, et al.
Published: (2024)
by: Yu, Seonghoon, et al.
Published: (2024)
VISAGE: Video Instance Segmentation with Appearance-Guided Enhancement
by: Kim, Hanjung, et al.
Published: (2023)
by: Kim, Hanjung, et al.
Published: (2023)
PHUMA: Physically-Grounded Humanoid Locomotion Dataset
by: Lee, Kyungmin, et al.
Published: (2025)
by: Lee, Kyungmin, et al.
Published: (2025)
Learning to Enhance Aperture Phasor Field for Non-Line-of-Sight Imaging
by: Cho, In, et al.
Published: (2024)
by: Cho, In, et al.
Published: (2024)
Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness
by: Lee, Hangyeol, et al.
Published: (2026)
by: Lee, Hangyeol, et al.
Published: (2026)
Video Token Merging for Long-form Video Understanding
by: Lee, Seon-Ho, et al.
Published: (2024)
by: Lee, Seon-Ho, et al.
Published: (2024)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
by: Yoon, Jieon, et al.
Published: (2026)
by: Yoon, Jieon, et al.
Published: (2026)
Enhancing Topological Dependencies in Spatio-Temporal Graphs with Cycle Message Passing Blocks
by: Lee, Minho, et al.
Published: (2024)
by: Lee, Minho, et al.
Published: (2024)
NegMerge: Sign-Consensual Weight Merging for Machine Unlearning
by: Kim, Hyo Seo, et al.
Published: (2024)
by: Kim, Hyo Seo, et al.
Published: (2024)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
by: Lee, Minhyun, et al.
Published: (2023)
by: Lee, Minhyun, et al.
Published: (2023)
Towards Scalable Handwriting Communication via EEG Decoding and Latent Embedding Integration
by: Kim, Jun-Young, et al.
Published: (2024)
by: Kim, Jun-Young, et al.
Published: (2024)
Motion-Oriented Compositional Neural Radiance Fields for Monocular Dynamic Human Modeling
by: Kim, Jaehyeok, et al.
Published: (2024)
by: Kim, Jaehyeok, et al.
Published: (2024)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
Domain Reduction Strategy for Non Line of Sight Imaging
by: Shim, Hyunbo, et al.
Published: (2023)
by: Shim, Hyunbo, et al.
Published: (2023)
Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
by: Kim, Sangmin, et al.
Published: (2026)
by: Kim, Sangmin, et al.
Published: (2026)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
VideoMamba: Spatio-Temporal Selective State Space Model
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Discrete Dictionary-based Decomposition Layer for Structured Representation Learning
by: Park, Taewon, et al.
Published: (2024)
by: Park, Taewon, et al.
Published: (2024)
ES-Merging: Biological MLLM Merging via Embedding Space Signals
by: Lee, Wonbin, et al.
Published: (2026)
by: Lee, Wonbin, et al.
Published: (2026)
Autoregressive Universal Video Segmentation Model
by: Heo, Miran, et al.
Published: (2025)
by: Heo, Miran, et al.
Published: (2025)
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
by: Kim, Seeyeon, et al.
Published: (2026)
by: Kim, Seeyeon, et al.
Published: (2026)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
by: Kim, Donghu, et al.
Published: (2024)
by: Kim, Donghu, et al.
Published: (2024)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
Transport Dynamics of Water Molecules Confined between Lipid Membranes
by: Lee, Minho, et al.
Published: (2024)
by: Lee, Minho, et al.
Published: (2024)
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
by: Kim, Hyunseung, et al.
Published: (2024)
by: Kim, Hyunseung, et al.
Published: (2024)
SyMerge: From Non-Interference to Synergistic Merging via Single-Layer Adaptation
by: Jung, Aecheon, et al.
Published: (2024)
by: Jung, Aecheon, et al.
Published: (2024)
Optimized Layerwise Approximation for Efficient Private Inference on Fully Homomorphic Encryption
by: Lee, Junghyun, et al.
Published: (2023)
by: Lee, Junghyun, et al.
Published: (2023)
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025)
by: Lee, Min-Jeong, et al.
Published: (2025)
Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos
by: Choi, Changwoon, et al.
Published: (2024)
by: Choi, Changwoon, et al.
Published: (2024)
AutiHero: Engaging Parents in Creating Personalized, Multi-path Social Narratives for Autistic Children
by: Lee, Jungeun, et al.
Published: (2025)
by: Lee, Jungeun, et al.
Published: (2025)
Similar Items
-
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025) -
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
by: Ahn, Geo, et al.
Published: (2026) -
Classification Matters: Improving Video Action Detection with Class-Specific Attention
by: Lee, Jinsung, et al.
Published: (2024) -
Video-Oasis: Rethinking Evaluation of Video Understanding
by: Lim, Geuntaek, et al.
Published: (2026) -
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)