SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khaki, Samir, Guo, Junxian, Tang, Jiaming, Yang, Shang, Chen, Yukang, Plataniotis, Konstantinos N., Lu, Yao, Han, Song, Liu, Zhijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
von: Khaki, Samir, et al.
Veröffentlicht: (2025)
The Need for Speed: Pruning Transformers with One Recipe
von: Khaki, Samir, et al.
Veröffentlicht: (2024)
von: Khaki, Samir, et al.
Veröffentlicht: (2024)
ProbMCL: Simple Probabilistic Contrastive Learning for Multi-label Visual Classification
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2024)
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2024)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
von: Liu, Zhijian, et al.
Veröffentlicht: (2024)
von: Liu, Zhijian, et al.
Veröffentlicht: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)
von: Yang, Shang, et al.
Veröffentlicht: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
ATOM: Attention Mixer for Efficient Dataset Distillation
von: Khaki, Samir, et al.
Veröffentlicht: (2024)
von: Khaki, Samir, et al.
Veröffentlicht: (2024)
Data-to-Model Distillation: Data-Efficient Learning Framework
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2024)
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2024)
DataDAM: Efficient Dataset Distillation with Attention Matching
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2023)
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2023)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
Active Inference and Reinforcement Learning: A unified inference on continuous state and action spaces under partial observability
von: Malekzadeh, Parvin, et al.
Veröffentlicht: (2022)
von: Malekzadeh, Parvin, et al.
Veröffentlicht: (2022)
VILA$^2$: VILA Augmented VILA
von: Fang, Yunhao, et al.
Veröffentlicht: (2024)
von: Fang, Yunhao, et al.
Veröffentlicht: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
VILA: On Pre-training for Visual Language Models
von: Lin, Ji, et al.
Veröffentlicht: (2023)
von: Lin, Ji, et al.
Veröffentlicht: (2023)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2025)
von: Tang, Jiaming, et al.
Veröffentlicht: (2025)
Emphasizing Discriminative Features for Dataset Distillation in Complex Scenarios
von: Wang, Kai, et al.
Veröffentlicht: (2024)
von: Wang, Kai, et al.
Veröffentlicht: (2024)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
von: Wu, Yecheng, et al.
Veröffentlicht: (2024)
von: Wu, Yecheng, et al.
Veröffentlicht: (2024)
Prioritize Alignment in Dataset Distillation
von: Li, Zekai, et al.
Veröffentlicht: (2024)
von: Li, Zekai, et al.
Veröffentlicht: (2024)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
EGM: Efficient Visual Grounding Language Models
von: Zhan, Guanqi, et al.
Veröffentlicht: (2026)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2026)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
XAttention: Block Sparse Attention with Antidiagonal Scoring
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
GSTAM: Efficient Graph Distillation with Structural Attention-Matching
von: Rasti-Meymandi, Arash, et al.
Veröffentlicht: (2024)
von: Rasti-Meymandi, Arash, et al.
Veröffentlicht: (2024)
A unified uncertainty-aware exploration: Combining epistemic and aleatory uncertainty
von: Malekzadeh, Parvin, et al.
Veröffentlicht: (2024)
von: Malekzadeh, Parvin, et al.
Veröffentlicht: (2024)
Uncertainty-aware transfer across tasks using hybrid model-based successor feature reinforcement learning
von: Malekzadeh, Parvin, et al.
Veröffentlicht: (2023)
von: Malekzadeh, Parvin, et al.
Veröffentlicht: (2023)
Dynamic Vision Mamba
von: Wu, Mengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Mengxuan, et al.
Veröffentlicht: (2025)
PanFlow: Decoupled Motion Control for Panoramic Video Generation
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
Activity Sparsity Complements Weight Sparsity for Efficient RNN Inference
von: Mukherji, Rishav, et al.
Veröffentlicht: (2023)
von: Mukherji, Rishav, et al.
Veröffentlicht: (2023)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
von: Tang, Haotian, et al.
Veröffentlicht: (2024)
von: Tang, Haotian, et al.
Veröffentlicht: (2024)
Sparsity-Constraint Optimization via Splicing Iteration
von: Zhu, Jin, et al.
Veröffentlicht: (2024)
von: Zhu, Jin, et al.
Veröffentlicht: (2024)
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
von: Lin, Chaofan, et al.
Veröffentlicht: (2025)
IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
von: Li, Junxian, et al.
Veröffentlicht: (2025)
von: Li, Junxian, et al.
Veröffentlicht: (2025)
VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge
von: Nath, Vishwesh, et al.
Veröffentlicht: (2024)
von: Nath, Vishwesh, et al.
Veröffentlicht: (2024)
An examination of the hierarchy problem beyond the Standard Model
von: Khaki, Seyed
Veröffentlicht: (2023)
von: Khaki, Seyed
Veröffentlicht: (2023)
Original $\mathbb F_1$ in emergent spacetime
von: Khaki, Seyed
Veröffentlicht: (2024)
von: Khaki, Seyed
Veröffentlicht: (2024)
Ähnliche Einträge
-
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
von: Khaki, Samir, et al.
Veröffentlicht: (2025) -
The Need for Speed: Pruning Transformers with One Recipe
von: Khaki, Samir, et al.
Veröffentlicht: (2024) -
ProbMCL: Simple Probabilistic Contrastive Learning for Multi-label Visual Classification
von: Sajedi, Ahmad, et al.
Veröffentlicht: (2024) -
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
von: Liu, Zhijian, et al.
Veröffentlicht: (2024) -
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)