HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914203390967808 |
|---|---|
| author | HyperAI Team Liu, Yuchen Han, Kaiyang Xia, Zhiqiang Dong, Yuhang Song, Chen Tang, Kangyu Xu, Jiaming Feng, Xiushi Yu, WenXuan Peng, Li Wang, Mingyang Wang, Kai Yang, Changpeng Li, Yang Lu, Haoyu Wang, Hao Xu, Bingna Liu, Guangyao Huang, Long Guo, Kaibin Wu, Jinyang Wu, Dan Wang, Hongzhen Zhou, Peng Nie, Shuai Wang, Shande Shi, Runyu Huang, Ying |
| author_facet | HyperAI Team Liu, Yuchen Han, Kaiyang Xia, Zhiqiang Dong, Yuhang Song, Chen Tang, Kangyu Xu, Jiaming Feng, Xiushi Yu, WenXuan Peng, Li Wang, Mingyang Wang, Kai Yang, Changpeng Li, Yang Lu, Haoyu Wang, Hao Xu, Bingna Liu, Guangyao Huang, Long Guo, Kaibin Wu, Jinyang Wu, Dan Wang, Hongzhen Zhou, Peng Nie, Shuai Wang, Shande Shi, Runyu Huang, Ying |
| contents | Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy directly on on-device environments. While small-parameter models are progressively endowed with strong general capabilities, standard Vision Transformer (ViT) encoders remain a critical bottleneck, suffering from excessive latency and memory consumption when processing high-resolution inputs.To address these challenges, we introduce HyperVL, an efficient multimodal large language model tailored for on-device inference. HyperVL adopts an image-tiling strategy to cap peak memory usage and incorporates two novel techniques: (1) a Visual Resolution Compressor (VRC) that adaptively predicts optimal encoding resolutions to eliminate redundant computation, and (2) Dual Consistency Learning (DCL), which aligns multi-scale ViT encoders within a unified framework, enabling dynamic switching between visual branches under a shared LLM. Extensive experiments demonstrate that HyperVL achieves state-of-the-art performance among models of comparable size across multiple benchmarks. Furthermore, it significantly significantly reduces latency and power consumption on real mobile devices, demonstrating its practicality for on-device multimodal inference. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_14052 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices HyperAI Team Liu, Yuchen Han, Kaiyang Xia, Zhiqiang Dong, Yuhang Song, Chen Tang, Kangyu Xu, Jiaming Feng, Xiushi Yu, WenXuan Peng, Li Wang, Mingyang Wang, Kai Yang, Changpeng Li, Yang Lu, Haoyu Wang, Hao Xu, Bingna Liu, Guangyao Huang, Long Guo, Kaibin Wu, Jinyang Wu, Dan Wang, Hongzhen Zhou, Peng Nie, Shuai Wang, Shande Shi, Runyu Huang, Ying Computer Vision and Pattern Recognition Computation and Language Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy directly on on-device environments. While small-parameter models are progressively endowed with strong general capabilities, standard Vision Transformer (ViT) encoders remain a critical bottleneck, suffering from excessive latency and memory consumption when processing high-resolution inputs.To address these challenges, we introduce HyperVL, an efficient multimodal large language model tailored for on-device inference. HyperVL adopts an image-tiling strategy to cap peak memory usage and incorporates two novel techniques: (1) a Visual Resolution Compressor (VRC) that adaptively predicts optimal encoding resolutions to eliminate redundant computation, and (2) Dual Consistency Learning (DCL), which aligns multi-scale ViT encoders within a unified framework, enabling dynamic switching between visual branches under a shared LLM. Extensive experiments demonstrate that HyperVL achieves state-of-the-art performance among models of comparable size across multiple benchmarks. Furthermore, it significantly significantly reduces latency and power consumption on real mobile devices, demonstrating its practicality for on-device multimodal inference. |
| title | HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2512.14052 |