VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Kichang, Kim, Seonjun, Kim, Minjae, Zhang, Nairan, Zhang, Chi, Lee, Youngki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911283450740736
author Yang, Kichang
Kim, Seonjun
Kim, Minjae
Zhang, Nairan
Zhang, Chi
Lee, Youngki
author_facet Yang, Kichang
Kim, Seonjun
Kim, Minjae
Zhang, Nairan
Zhang, Chi
Lee, Youngki
contents Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However, conventional sparsification remains model-centric, selecting neurons solely by activation magnitude and neglecting how access patterns influence flash performance. We present Neuron Chunking, an I/O-efficient sparsification strategy that operates on chunks (i.e., groups of contiguous neurons in memory) and couples neuron importance with storage access cost. The method models I/O latency through a lightweight abstraction of access contiguity and selects chunks with high utility, defined as neuron importance normalized by estimated latency. By aligning sparsification decisions with the underlying storage behavior, Neuron Chunking improves I/O efficiency by up to 4.65x and 5.76x on Jetson Orin Nano and Jetson AGX Orin, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18692
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
Yang, Kichang
Kim, Seonjun
Kim, Minjae
Zhang, Nairan
Zhang, Chi
Lee, Youngki
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Performance
Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However, conventional sparsification remains model-centric, selecting neurons solely by activation magnitude and neglecting how access patterns influence flash performance. We present Neuron Chunking, an I/O-efficient sparsification strategy that operates on chunks (i.e., groups of contiguous neurons in memory) and couples neuron importance with storage access cost. The method models I/O latency through a lightweight abstraction of access contiguity and selects chunks with high utility, defined as neuron importance normalized by estimated latency. By aligning sparsification decisions with the underlying storage behavior, Neuron Chunking improves I/O efficiency by up to 4.65x and 5.76x on Jetson Orin Nano and Jetson AGX Orin, respectively.
title VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Performance
url https://arxiv.org/abs/2511.18692