Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Xixi, Yang, Chen, Zhang, Dong, Dong, Pingcheng, Yang, Xin, Cheng, Kwang-Ting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory Efficient Transformer Adapter for Dense Predictions
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
Towards Customized Knowledge Distillation for Chip-Level Dense Image Predictions
by: Zhang, Dong, et al.
Published: (2024)
by: Zhang, Dong, et al.
Published: (2024)
Labeled-to-Unlabeled Distribution Alignment for Partially-Supervised Multi-Organ Medical Image Segmentation
by: Jiang, Xixi, et al.
Published: (2024)
by: Jiang, Xixi, et al.
Published: (2024)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
Video Token Merging for Long-form Video Understanding
by: Lee, Seon-Ho, et al.
Published: (2024)
by: Lee, Seon-Ho, et al.
Published: (2024)
BoNuS: Boundary Mining for Nuclei Segmentation with Partial Point Labels
by: Lin, Yi, et al.
Published: (2024)
by: Lin, Yi, et al.
Published: (2024)
Generalized Task-Driven Medical Image Quality Enhancement with Gradient Promotion
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)
by: Zhang, Luyuan, et al.
Published: (2026)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)
by: Liu, Shih-yang, et al.
Published: (2023)
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
by: Cheng, Zheng, et al.
Published: (2024)
by: Cheng, Zheng, et al.
Published: (2024)
ToSA: Token Merging with Spatial Awareness
by: Huang, Hsiang-Wei, et al.
Published: (2025)
by: Huang, Hsiang-Wei, et al.
Published: (2025)
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
by: Fan, Ziyang, et al.
Published: (2026)
by: Fan, Ziyang, et al.
Published: (2026)
Lossless Token Merging Even Without Fine-Tuning in Vision Transformers
by: Lee, Jaeyeon, et al.
Published: (2025)
by: Lee, Jaeyeon, et al.
Published: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
by: Lee, Dong Hoon, et al.
Published: (2024)
by: Lee, Dong Hoon, et al.
Published: (2024)
Sequential Token Merging: Revisiting Hidden States
by: Wen, Yan, et al.
Published: (2025)
by: Wen, Yan, et al.
Published: (2025)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
Video, How Do Your Tokens Merge?
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets
by: Liu, Zhixuan, et al.
Published: (2026)
by: Liu, Zhixuan, et al.
Published: (2026)
HoliTom: Holistic Token Merging for Fast Video Large Language Models
by: Shao, Kele, et al.
Published: (2025)
by: Shao, Kele, et al.
Published: (2025)
Instrument-tissue Interaction Detection Framework for Surgical Video Understanding
by: Lin, Wenjun, et al.
Published: (2024)
by: Lin, Wenjun, et al.
Published: (2024)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
by: Shen, Leqi, et al.
Published: (2024)
by: Shen, Leqi, et al.
Published: (2024)
Aligning Medical Images with General Knowledge from Large Language Models
by: Fang, Xiao, et al.
Published: (2024)
by: Fang, Xiao, et al.
Published: (2024)
ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack
by: Gao, Ziyi, et al.
Published: (2024)
by: Gao, Ziyi, et al.
Published: (2024)
Layer-Aware Video Composition via Split-then-Merge
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025)
by: Zhang, Siyang, et al.
Published: (2025)
Importance-Based Token Merging for Efficient Image and Video Generation
by: Wu, Haoyu, et al.
Published: (2024)
by: Wu, Haoyu, et al.
Published: (2024)
Exploring Enhanced Contextual Information for Video-Level Object Tracking
by: Kang, Ben, et al.
Published: (2024)
by: Kang, Ben, et al.
Published: (2024)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
by: Zhang, Liang, et al.
Published: (2024)
by: Zhang, Liang, et al.
Published: (2024)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
by: Liu, Xinxin, et al.
Published: (2026)
by: Liu, Xinxin, et al.
Published: (2026)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
by: Jin, Hongbo, et al.
Published: (2025)
by: Jin, Hongbo, et al.
Published: (2025)
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
by: Wu, Jinlin, et al.
Published: (2026)
by: Wu, Jinlin, et al.
Published: (2026)
EarlyTom: Early Token Compression Completes Fast Video Understanding
by: Wang, Hesong, et al.
Published: (2026)
by: Wang, Hesong, et al.
Published: (2026)
Surgical Video Understanding with Label Interpolation
by: Kim, Garam, et al.
Published: (2025)
by: Kim, Garam, et al.
Published: (2025)
MReg: A Novel Regression Model with MoE-based Video Feature Mining for Mitral Regurgitation Diagnosis
by: Liu, Zhe, et al.
Published: (2025)
by: Liu, Zhe, et al.
Published: (2025)
Realistic Surgical Simulation from Monocular Videos
by: Wang, Kailing, et al.
Published: (2024)
by: Wang, Kailing, et al.
Published: (2024)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
by: Guo, Yanan, et al.
Published: (2025)
by: Guo, Yanan, et al.
Published: (2025)
Efficient Visual Transformer by Learnable Token Merging
by: Wang, Yancheng, et al.
Published: (2024)
by: Wang, Yancheng, et al.
Published: (2024)
Similar Items
-
Memory Efficient Transformer Adapter for Dense Predictions
by: Zhang, Dong, et al.
Published: (2025) -
Towards Customized Knowledge Distillation for Chip-Level Dense Image Predictions
by: Zhang, Dong, et al.
Published: (2024) -
Labeled-to-Unlabeled Distribution Alignment for Partially-Supervised Multi-Organ Medical Image Segmentation
by: Jiang, Xixi, et al.
Published: (2024) -
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023) -
Video Token Merging for Long-form Video Understanding
by: Lee, Seon-Ho, et al.
Published: (2024)