Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915685795364864 |
|---|---|
| author | Panchal, Kunjal Mitra, Saayan Sarkhel, Somdeb Wang, Haoliang Dasgupta, Ishita Wu, Gang Guan, Hui |
| author_facet | Panchal, Kunjal Mitra, Saayan Sarkhel, Somdeb Wang, Haoliang Dasgupta, Ishita Wu, Gang Guan, Hui |
| contents | Recent advances in video-language models have enabled powerful applications like video retrieval, captioning, and assembly. However, executing such multi-stage pipelines efficiently on mobile devices remains challenging due to redundant model loads and fragmented execution. We introduce Atom, an on-device system that restructures video-language pipelines for fast and efficient execution. Atom decomposes a billion-parameter model into reusable modules, such as the visual encoder and language decoder, and reuses them across subtasks like captioning, reasoning, and indexing. This reuse-centric design eliminates repeated model loading and enables parallel execution, reducing end-to-end latency without sacrificing performance. On commodity smartphones, Atom achieves 27--33% faster execution compared to non-reuse baselines, with only marginal performance drop ($\leq$ 2.3 Recall@1 in retrieval, $\leq$ 1.5 CIDEr in captioning). These results position Atom as a practical, scalable approach for efficient video-language understanding on edge devices. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_17108 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse Panchal, Kunjal Mitra, Saayan Sarkhel, Somdeb Wang, Haoliang Dasgupta, Ishita Wu, Gang Guan, Hui Machine Learning Multimedia Recent advances in video-language models have enabled powerful applications like video retrieval, captioning, and assembly. However, executing such multi-stage pipelines efficiently on mobile devices remains challenging due to redundant model loads and fragmented execution. We introduce Atom, an on-device system that restructures video-language pipelines for fast and efficient execution. Atom decomposes a billion-parameter model into reusable modules, such as the visual encoder and language decoder, and reuses them across subtasks like captioning, reasoning, and indexing. This reuse-centric design eliminates repeated model loading and enables parallel execution, reducing end-to-end latency without sacrificing performance. On commodity smartphones, Atom achieves 27--33% faster execution compared to non-reuse baselines, with only marginal performance drop ($\leq$ 2.3 Recall@1 in retrieval, $\leq$ 1.5 CIDEr in captioning). These results position Atom as a practical, scalable approach for efficient video-language understanding on edge devices. |
| title | Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse |
| topic | Machine Learning Multimedia |
| url | https://arxiv.org/abs/2512.17108 |