DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Aditya Kumar, Kandala, Hitesh, Brahma, Pratik Prabhanjan, Liu, Zicheng, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
von: Hu, Junshan, et al.
Veröffentlicht: (2025)
von: Hu, Junshan, et al.
Veröffentlicht: (2025)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
von: Joshi, Vinay, et al.
Veröffentlicht: (2025)
von: Joshi, Vinay, et al.
Veröffentlicht: (2025)
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
von: Li, Geng, et al.
Veröffentlicht: (2026)
von: Li, Geng, et al.
Veröffentlicht: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
DocVLM: Make Your VLM an Efficient Reader
von: Nacson, Mor Shpigel, et al.
Veröffentlicht: (2024)
von: Nacson, Mor Shpigel, et al.
Veröffentlicht: (2024)
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
von: Qian, Chen, et al.
Veröffentlicht: (2026)
von: Qian, Chen, et al.
Veröffentlicht: (2026)
Similarity-Aware Token Pruning: Your VLM but Faster
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
InspectVLM: Unified in Theory, Unreliable in Practice
von: Wallace, Conor, et al.
Veröffentlicht: (2025)
von: Wallace, Conor, et al.
Veröffentlicht: (2025)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
von: Chen, Tao, et al.
Veröffentlicht: (2025)
von: Chen, Tao, et al.
Veröffentlicht: (2025)
SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
EgoVLM: Policy Optimization for Egocentric Video Understanding
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
von: Vinod, Ashwin, et al.
Veröffentlicht: (2025)
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
von: Xie, Rongchang, et al.
Veröffentlicht: (2024)
von: Xie, Rongchang, et al.
Veröffentlicht: (2024)
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
von: Tai, Yan, et al.
Veröffentlicht: (2025)
von: Tai, Yan, et al.
Veröffentlicht: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Pix2Gif: Motion-Guided Diffusion for GIF Generation
von: Kandala, Hitesh, et al.
Veröffentlicht: (2024)
von: Kandala, Hitesh, et al.
Veröffentlicht: (2024)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
von: Xue, Xizhe, et al.
Veröffentlicht: (2024)
von: Xue, Xizhe, et al.
Veröffentlicht: (2024)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
DiffSparse: Accelerating Diffusion Transformers with Learned Token Sparsity
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
von: Zhu, Haowei, et al.
Veröffentlicht: (2026)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
Enhancing Video Transformers for Action Understanding with VLM-aided Training
von: Lu, Hui, et al.
Veröffentlicht: (2024)
von: Lu, Hui, et al.
Veröffentlicht: (2024)
UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
von: Bai, Chengyu, et al.
Veröffentlicht: (2025)
von: Bai, Chengyu, et al.
Veröffentlicht: (2025)
EffiMiniVLM: A Compact Dual-Encoder Regression Framework
von: Khor, Yin-Loon, et al.
Veröffentlicht: (2026)
von: Khor, Yin-Loon, et al.
Veröffentlicht: (2026)
Rethinking VLM Representation for VLA Initialization
von: Lin, Weifeng, et al.
Veröffentlicht: (2026)
von: Lin, Weifeng, et al.
Veröffentlicht: (2026)
MOVi: Training-free Text-conditioned Multi-Object Video Generation
von: Rahman, Aimon, et al.
Veröffentlicht: (2025)
von: Rahman, Aimon, et al.
Veröffentlicht: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
von: Zhang, Di, et al.
Veröffentlicht: (2024)
von: Zhang, Di, et al.
Veröffentlicht: (2024)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
LADDER: An Efficient Framework for Video Frame Interpolation
von: Shen, Tong, et al.
Veröffentlicht: (2024)
von: Shen, Tong, et al.
Veröffentlicht: (2024)
Learning from Online Videos at Inference Time for Computer-Use Agents
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
von: Singh, Shivam, et al.
Veröffentlicht: (2026) -
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
von: Hu, Junshan, et al.
Veröffentlicht: (2025) -
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
von: Joshi, Vinay, et al.
Veröffentlicht: (2025) -
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
von: Li, Geng, et al.
Veröffentlicht: (2026) -
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)