SpecVLM: Fast Speculative Decoding in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Haiduo, Yang, Fuwei, Liu, Zhenhua, Yin, Xuanwu, Li, Dong, Ren, Pengju, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
Partial Convolution Meets Visual Attention
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Nearly Lossless Adaptive Bit Switching
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
Partial Channel Network: Compute Fewer, Perform Better
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
KernelDNA: Dynamic Kernel Sharing via Decoupled Naive Adapters
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
GeGS-PCR: Effective and Robust 3D Point Cloud Registration with Two-Stage Color-Enhanced Geometric-3DGS Fusion
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
von: Shen, Hui, et al.
Veröffentlicht: (2026)
von: Shen, Hui, et al.
Veröffentlicht: (2026)
E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
von: Shen, Tong, et al.
Veröffentlicht: (2025)
von: Shen, Tong, et al.
Veröffentlicht: (2025)
MonoGS++: Fast and Accurate Monocular RGB Gaussian SLAM
von: Li, Renwu, et al.
Veröffentlicht: (2025)
von: Li, Renwu, et al.
Veröffentlicht: (2025)
FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Fast Occupancy Network
von: Lu, Mingjie, et al.
Veröffentlicht: (2024)
von: Lu, Mingjie, et al.
Veröffentlicht: (2024)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
FloorplanVLM: A Vision-Language Model for Floorplan Vectorization
von: Liu, Yuanqing, et al.
Veröffentlicht: (2026)
von: Liu, Yuanqing, et al.
Veröffentlicht: (2026)
Edit as You See: Image-guided Video Editing via Masked Motion Modeling
von: Huang, Zhi-Lin, et al.
Veröffentlicht: (2025)
von: Huang, Zhi-Lin, et al.
Veröffentlicht: (2025)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
von: Tong, Yujia, et al.
Veröffentlicht: (2026)
von: Tong, Yujia, et al.
Veröffentlicht: (2026)
AMD-Hummingbird: Towards an Efficient Text-to-Video Model
von: Isobe, Takashi, et al.
Veröffentlicht: (2025)
von: Isobe, Takashi, et al.
Veröffentlicht: (2025)
HSD: Training-Free Acceleration for Document Parsing Vision-Language Model with Hierarchical Speculative Decoding
von: Liao, Wenhui, et al.
Veröffentlicht: (2026)
von: Liao, Wenhui, et al.
Veröffentlicht: (2026)
SDXS: Real-Time One-Step Latent Diffusion Models with Image Conditions
von: Song, Yuda, et al.
Veröffentlicht: (2024)
von: Song, Yuda, et al.
Veröffentlicht: (2024)
FastVLM: Efficient Vision Encoding for Vision Language Models
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2024)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
von: Kong, Quan, et al.
Veröffentlicht: (2026)
von: Kong, Quan, et al.
Veröffentlicht: (2026)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2023)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2023)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
SpecDiff: Accelerating Diffusion Model Inference with Self-Speculation
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
LADDER: An Efficient Framework for Video Frame Interpolation
von: Shen, Tong, et al.
Veröffentlicht: (2024)
von: Shen, Tong, et al.
Veröffentlicht: (2024)
Hierarchical Cross-modal Prompt Learning for Vision-Language Models
von: Zheng, Hao, et al.
Veröffentlicht: (2025)
von: Zheng, Hao, et al.
Veröffentlicht: (2025)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
U-VLM: Hierarchical Vision Language Modeling for Report Generation
von: Shi, Pengcheng, et al.
Veröffentlicht: (2026)
von: Shi, Pengcheng, et al.
Veröffentlicht: (2026)
Slot-VLM: SlowFast Slots for Video-Language Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
von: Li, Juncheng, et al.
Veröffentlicht: (2025)
von: Li, Juncheng, et al.
Veröffentlicht: (2025)
Speculative Decoding Reimagined for Multimodal Large Language Models
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
von: Huang, Zhipeng, et al.
Veröffentlicht: (2024)
von: Huang, Zhipeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025) -
Partial Convolution Meets Visual Attention
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025) -
Nearly Lossless Adaptive Bit Switching
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)