Cached Adaptive Token Merging: Dynamic Token Reduction and Redundant Computation Elimination in Diffusion Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Saghatchian, Omid, Moghadam, Atiyeh Gh., Nickabadi, Ahmad |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
par: Lee, Yuna, et autres
Publié: (2026)
par: Lee, Yuna, et autres
Publié: (2026)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
par: Kim, Kwonyoung, et autres
Publié: (2025)
par: Kim, Kwonyoung, et autres
Publié: (2025)
LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
par: Hu, Lianyu, et autres
Publié: (2025)
par: Hu, Lianyu, et autres
Publié: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
par: Zhang, Luyuan, et autres
Publié: (2026)
par: Zhang, Luyuan, et autres
Publié: (2026)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
par: Shang, Yuzhang, et autres
Publié: (2024)
par: Shang, Yuzhang, et autres
Publié: (2024)
Token Caching for Diffusion Transformer Acceleration
par: Lou, Jinming, et autres
Publié: (2024)
par: Lou, Jinming, et autres
Publié: (2024)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
par: Shu, Zhijian, et autres
Publié: (2025)
par: Shu, Zhijian, et autres
Publié: (2025)
Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
par: Zheng, Zhixin, et autres
Publié: (2025)
par: Zheng, Zhixin, et autres
Publié: (2025)
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
par: Li, Duo, et autres
Publié: (2025)
par: Li, Duo, et autres
Publié: (2025)
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
par: Feng, Weilun, et autres
Publié: (2026)
par: Feng, Weilun, et autres
Publié: (2026)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
par: Kilian, Maciej, et autres
Publié: (2024)
par: Kilian, Maciej, et autres
Publié: (2024)
DORA: Dynamic Online Reinforcement Agent for Token Merging in Vision Transformers
par: He, Kaixuan, et autres
Publié: (2026)
par: He, Kaixuan, et autres
Publié: (2026)
Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness
par: Lee, Hangyeol, et autres
Publié: (2026)
par: Lee, Hangyeol, et autres
Publié: (2026)
Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
par: Cheng, Tianle, et autres
Publié: (2025)
par: Cheng, Tianle, et autres
Publié: (2025)
ToSA: Token Merging with Spatial Awareness
par: Huang, Hsiang-Wei, et autres
Publié: (2025)
par: Huang, Hsiang-Wei, et autres
Publié: (2025)
Video, How Do Your Tokens Merge?
par: Pollard, Sam, et autres
Publié: (2025)
par: Pollard, Sam, et autres
Publié: (2025)
Sequential Token Merging: Revisiting Hidden States
par: Wen, Yan, et autres
Publié: (2025)
par: Wen, Yan, et autres
Publié: (2025)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
par: Li, Hongliang, et autres
Publié: (2025)
par: Li, Hongliang, et autres
Publié: (2025)
Towards Efficient Vision State Space Models via Token Merging
par: Park, Jinyoung, et autres
Publié: (2025)
par: Park, Jinyoung, et autres
Publié: (2025)
Video Token Merging for Long-form Video Understanding
par: Lee, Seon-Ho, et autres
Publié: (2024)
par: Lee, Seon-Ho, et autres
Publié: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
par: Norouzi, Narges, et autres
Publié: (2024)
par: Norouzi, Narges, et autres
Publié: (2024)
TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
par: Zhang, Hao, et autres
Publié: (2025)
par: Zhang, Hao, et autres
Publié: (2025)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
par: Xu, Siyu, et autres
Publié: (2025)
par: Xu, Siyu, et autres
Publié: (2025)
FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
par: Mehraban, Soroush, et autres
Publié: (2025)
par: Mehraban, Soroush, et autres
Publié: (2025)
ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack
par: Gao, Ziyi, et autres
Publié: (2024)
par: Gao, Ziyi, et autres
Publié: (2024)
Accelerating Diffusion Transformers with Token-wise Feature Caching
par: Zou, Chang, et autres
Publié: (2024)
par: Zou, Chang, et autres
Publié: (2024)
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
par: Yin, Shicheng, et autres
Publié: (2025)
par: Yin, Shicheng, et autres
Publié: (2025)
Guiding Token-Sparse Diffusion Models
par: Krause, Felix, et autres
Publié: (2026)
par: Krause, Felix, et autres
Publié: (2026)
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
par: Chang, Shuochen, et autres
Publié: (2025)
par: Chang, Shuochen, et autres
Publié: (2025)
Token Bottleneck: One Token to Remember Dynamics
par: Kim, Taekyung, et autres
Publié: (2025)
par: Kim, Taekyung, et autres
Publié: (2025)
HoliTom: Holistic Token Merging for Fast Video Large Language Models
par: Shao, Kele, et autres
Publié: (2025)
par: Shao, Kele, et autres
Publié: (2025)
FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning
par: Guo, Hang, et autres
Publié: (2025)
par: Guo, Hang, et autres
Publié: (2025)
Dynamic Token Reduction during Generation for Vision Language Models
par: Liang, Xiaoyu, et autres
Publié: (2025)
par: Liang, Xiaoyu, et autres
Publié: (2025)
HTTM: Head-wise Temporal Token Merging for Faster VGGT
par: Wang, Weitian, et autres
Publié: (2025)
par: Wang, Weitian, et autres
Publié: (2025)
Importance-Based Token Merging for Efficient Image and Video Generation
par: Wu, Haoyu, et autres
Publié: (2024)
par: Wu, Haoyu, et autres
Publié: (2024)
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
par: Zhang, Yiming, et autres
Publié: (2024)
par: Zhang, Yiming, et autres
Publié: (2024)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
par: Zhang, Kejia, et autres
Publié: (2025)
par: Zhang, Kejia, et autres
Publié: (2025)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
par: Bagrov, Natan, et autres
Publié: (2025)
par: Bagrov, Natan, et autres
Publié: (2025)
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
par: Yao, Linli, et autres
Publié: (2025)
par: Yao, Linli, et autres
Publié: (2025)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
par: Wang, Zirui, et autres
Publié: (2023)
par: Wang, Zirui, et autres
Publié: (2023)
Documents similaires
-
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
par: Lee, Yuna, et autres
Publié: (2026) -
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
par: Kim, Kwonyoung, et autres
Publié: (2025) -
LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
par: Hu, Lianyu, et autres
Publié: (2025) -
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
par: Zhang, Luyuan, et autres
Publié: (2026) -
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
par: Shang, Yuzhang, et autres
Publié: (2024)