HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: HyperAI Team, Liu, Yuchen, Han, Kaiyang, Xia, Zhiqiang, Dong, Yuhang, Song, Chen, Tang, Kangyu, Xu, Jiaming, Feng, Xiushi, Yu, WenXuan, Peng, Li, Wang, Mingyang, Wang, Kai, Yang, Changpeng, Li, Yang, Lu, Haoyu, Wang, Hao, Xu, Bingna, Liu, Guangyao, Huang, Long, Guo, Kaibin, Wu, Jinyang, Wu, Dan, Wang, Hongzhen, Zhou, Peng, Nie, Shuai, Wang, Shande, Shi, Runyu, Huang, Ying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914203390967808
author HyperAI Team
Liu, Yuchen
Han, Kaiyang
Xia, Zhiqiang
Dong, Yuhang
Song, Chen
Tang, Kangyu
Xu, Jiaming
Feng, Xiushi
Yu, WenXuan
Peng, Li
Wang, Mingyang
Wang, Kai
Yang, Changpeng
Li, Yang
Lu, Haoyu
Wang, Hao
Xu, Bingna
Liu, Guangyao
Huang, Long
Guo, Kaibin
Wu, Jinyang
Wu, Dan
Wang, Hongzhen
Zhou, Peng
Nie, Shuai
Wang, Shande
Shi, Runyu
Huang, Ying
author_facet HyperAI Team
Liu, Yuchen
Han, Kaiyang
Xia, Zhiqiang
Dong, Yuhang
Song, Chen
Tang, Kangyu
Xu, Jiaming
Feng, Xiushi
Yu, WenXuan
Peng, Li
Wang, Mingyang
Wang, Kai
Yang, Changpeng
Li, Yang
Lu, Haoyu
Wang, Hao
Xu, Bingna
Liu, Guangyao
Huang, Long
Guo, Kaibin
Wu, Jinyang
Wu, Dan
Wang, Hongzhen
Zhou, Peng
Nie, Shuai
Wang, Shande
Shi, Runyu
Huang, Ying
contents Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy directly on on-device environments. While small-parameter models are progressively endowed with strong general capabilities, standard Vision Transformer (ViT) encoders remain a critical bottleneck, suffering from excessive latency and memory consumption when processing high-resolution inputs.To address these challenges, we introduce HyperVL, an efficient multimodal large language model tailored for on-device inference. HyperVL adopts an image-tiling strategy to cap peak memory usage and incorporates two novel techniques: (1) a Visual Resolution Compressor (VRC) that adaptively predicts optimal encoding resolutions to eliminate redundant computation, and (2) Dual Consistency Learning (DCL), which aligns multi-scale ViT encoders within a unified framework, enabling dynamic switching between visual branches under a shared LLM. Extensive experiments demonstrate that HyperVL achieves state-of-the-art performance among models of comparable size across multiple benchmarks. Furthermore, it significantly significantly reduces latency and power consumption on real mobile devices, demonstrating its practicality for on-device multimodal inference.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14052
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
HyperAI Team
Liu, Yuchen
Han, Kaiyang
Xia, Zhiqiang
Dong, Yuhang
Song, Chen
Tang, Kangyu
Xu, Jiaming
Feng, Xiushi
Yu, WenXuan
Peng, Li
Wang, Mingyang
Wang, Kai
Yang, Changpeng
Li, Yang
Lu, Haoyu
Wang, Hao
Xu, Bingna
Liu, Guangyao
Huang, Long
Guo, Kaibin
Wu, Jinyang
Wu, Dan
Wang, Hongzhen
Zhou, Peng
Nie, Shuai
Wang, Shande
Shi, Runyu
Huang, Ying
Computer Vision and Pattern Recognition
Computation and Language
Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy directly on on-device environments. While small-parameter models are progressively endowed with strong general capabilities, standard Vision Transformer (ViT) encoders remain a critical bottleneck, suffering from excessive latency and memory consumption when processing high-resolution inputs.To address these challenges, we introduce HyperVL, an efficient multimodal large language model tailored for on-device inference. HyperVL adopts an image-tiling strategy to cap peak memory usage and incorporates two novel techniques: (1) a Visual Resolution Compressor (VRC) that adaptively predicts optimal encoding resolutions to eliminate redundant computation, and (2) Dual Consistency Learning (DCL), which aligns multi-scale ViT encoders within a unified framework, enabling dynamic switching between visual branches under a shared LLM. Extensive experiments demonstrate that HyperVL achieves state-of-the-art performance among models of comparable size across multiple benchmarks. Furthermore, it significantly significantly reduces latency and power consumption on real mobile devices, demonstrating its practicality for on-device multimodal inference.
title HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2512.14052