INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Šabanović, Ahmed, Maliakel, Paul Joe, Brandić, Ivona
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917509640224768
author Šabanović, Ahmed
Maliakel, Paul Joe
Brandić, Ivona
author_facet Šabanović, Ahmed
Maliakel, Paul Joe
Brandić, Ivona
contents Edge deployment of Vision-Language Models (VLMs) faces a tradeoff between latency and accuracy: cloud execution provides high-quality predictions but incurs communication delay and energy cost, while edge-only execution is faster but less accurate due to limited model capacity. This trade-off is further complicated by heterogeneity in image quality and reasoning complexity, making static placement suboptimal. We present INAR-VL, a lightweight edge-cloud routing system for multimodal inference in a two-tier deployment. INAR-VL maintains complementary VLMs across edge and cloud and uses lightweight image and text complexity signals to guide routing and model selection, executing simple queries locally while offloading complex ones when beneficial. Evaluation on visual question answering shows that INAR-VL executes 36% of requests on the edge, reduces latency by 24%, lowers energy by 26%, and preserves 97% of cloud-level accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18853
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
Šabanović, Ahmed
Maliakel, Paul Joe
Brandić, Ivona
Machine Learning
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
C.3; I.2.10; I.2.6; D.4.8
Edge deployment of Vision-Language Models (VLMs) faces a tradeoff between latency and accuracy: cloud execution provides high-quality predictions but incurs communication delay and energy cost, while edge-only execution is faster but less accurate due to limited model capacity. This trade-off is further complicated by heterogeneity in image quality and reasoning complexity, making static placement suboptimal. We present INAR-VL, a lightweight edge-cloud routing system for multimodal inference in a two-tier deployment. INAR-VL maintains complementary VLMs across edge and cloud and uses lightweight image and text complexity signals to guide routing and model selection, executing simple queries locally while offloading complex ones when beneficial. Evaluation on visual question answering shows that INAR-VL executes 36% of requests on the edge, reduces latency by 24%, lowers energy by 26%, and preserves 97% of cloud-level accuracy.
title INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
topic Machine Learning
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
C.3; I.2.10; I.2.6; D.4.8
url https://arxiv.org/abs/2605.18853