Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Panchal, Kunjal, Mitra, Saayan, Sarkhel, Somdeb, Wang, Haoliang, Dasgupta, Ishita, Wu, Gang, Guan, Hui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915685795364864
author Panchal, Kunjal
Mitra, Saayan
Sarkhel, Somdeb
Wang, Haoliang
Dasgupta, Ishita
Wu, Gang
Guan, Hui
author_facet Panchal, Kunjal
Mitra, Saayan
Sarkhel, Somdeb
Wang, Haoliang
Dasgupta, Ishita
Wu, Gang
Guan, Hui
contents Recent advances in video-language models have enabled powerful applications like video retrieval, captioning, and assembly. However, executing such multi-stage pipelines efficiently on mobile devices remains challenging due to redundant model loads and fragmented execution. We introduce Atom, an on-device system that restructures video-language pipelines for fast and efficient execution. Atom decomposes a billion-parameter model into reusable modules, such as the visual encoder and language decoder, and reuses them across subtasks like captioning, reasoning, and indexing. This reuse-centric design eliminates repeated model loading and enables parallel execution, reducing end-to-end latency without sacrificing performance. On commodity smartphones, Atom achieves 27--33% faster execution compared to non-reuse baselines, with only marginal performance drop ($\leq$ 2.3 Recall@1 in retrieval, $\leq$ 1.5 CIDEr in captioning). These results position Atom as a practical, scalable approach for efficient video-language understanding on edge devices.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17108
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
Panchal, Kunjal
Mitra, Saayan
Sarkhel, Somdeb
Wang, Haoliang
Dasgupta, Ishita
Wu, Gang
Guan, Hui
Machine Learning
Multimedia
Recent advances in video-language models have enabled powerful applications like video retrieval, captioning, and assembly. However, executing such multi-stage pipelines efficiently on mobile devices remains challenging due to redundant model loads and fragmented execution. We introduce Atom, an on-device system that restructures video-language pipelines for fast and efficient execution. Atom decomposes a billion-parameter model into reusable modules, such as the visual encoder and language decoder, and reuses them across subtasks like captioning, reasoning, and indexing. This reuse-centric design eliminates repeated model loading and enables parallel execution, reducing end-to-end latency without sacrificing performance. On commodity smartphones, Atom achieves 27--33% faster execution compared to non-reuse baselines, with only marginal performance drop ($\leq$ 2.3 Recall@1 in retrieval, $\leq$ 1.5 CIDEr in captioning). These results position Atom as a practical, scalable approach for efficient video-language understanding on edge devices.
title Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
topic Machine Learning
Multimedia
url https://arxiv.org/abs/2512.17108