AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Zhiwei, Song, Xiaohui, Wang, Nan, Liu, Yafei, Li, Chao, Li, Xin, Wang, Ruichen, Li, Zhihao, Qi, Qi, Cheng, Long, Hao, Dongze, Zheng, Quanlong, Zhang, Yanhao, Ji, Haobo, Ma, Jian, Zheng, Zhitong, Lin, Zhenyi, Deng, Haolin, Zou, Xin, Yin, Xiaojie, Wang, Ruilin, Cai, Liankai, Liu, Haijing, Qiu, Yuqing, Chen, Ke, Li, Zixian, Xie, Chi, Li, Huafei, Li, Chenxing, Wang, Chuangchuang, Tang, Kai, Zhu, Zhiguang, Gao, Wenmei, Wang, Rui, Wu, Jun, Liu, Chao, Xie, Qin, Chen, Chen, Lu, Haonan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915689469575168
author Jin, Zhiwei
Song, Xiaohui
Wang, Nan
Liu, Yafei
Li, Chao
Li, Xin
Wang, Ruichen
Li, Zhihao
Qi, Qi
Cheng, Long
Hao, Dongze
Zheng, Quanlong
Zhang, Yanhao
Ji, Haobo
Ma, Jian
Zheng, Zhitong
Lin, Zhenyi
Deng, Haolin
Zou, Xin
Yin, Xiaojie
Wang, Ruilin
Cai, Liankai
Liu, Haijing
Qiu, Yuqing
Chen, Ke
Li, Zixian
Xie, Chi
Li, Huafei
Li, Chenxing
Wang, Chuangchuang
Tang, Kai
Zhu, Zhiguang
Tang, Kai
Gao, Wenmei
Wang, Rui
Wu, Jun
Liu, Chao
Xie, Qin
Chen, Chen
Lu, Haonan
author_facet Jin, Zhiwei
Song, Xiaohui
Wang, Nan
Liu, Yafei
Li, Chao
Li, Xin
Wang, Ruichen
Li, Zhihao
Qi, Qi
Cheng, Long
Hao, Dongze
Zheng, Quanlong
Zhang, Yanhao
Ji, Haobo
Ma, Jian
Zheng, Zhitong
Lin, Zhenyi
Deng, Haolin
Zou, Xin
Yin, Xiaojie
Wang, Ruilin
Cai, Liankai
Liu, Haijing
Qiu, Yuqing
Chen, Ke
Li, Zixian
Xie, Chi
Li, Huafei
Li, Chenxing
Wang, Chuangchuang
Tang, Kai
Zhu, Zhiguang
Tang, Kai
Gao, Wenmei
Wang, Rui
Wu, Jun
Liu, Chao
Xie, Qin
Chen, Chen
Lu, Haonan
contents In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hundreds of billions of parameters, they significantly surpass the limitations in memory, power consumption, and computing capacity of edge devices such as mobile phones. This paper introduces AndesVL, a suite of mobile-side MLLMs with 0.6B to 4B parameters based on Qwen3's LLM and various visual encoders. We comprehensively outline the model architectures, training pipeline, and training data of AndesVL, which achieves first-tier performance across a wide range of open-source benchmarks, including fields such as text-rich image understanding, reasoning and math, multi-image comprehension, general VQA, hallucination mitigation, multilingual understanding, and GUI-related tasks when compared with state-of-the-art models of a similar scale. Furthermore, we introduce a 1+N LoRA architecture alongside a Quantization-Aware LoRA Fine-Tuning (QALFT) framework to facilitate efficient task adaptation and model compression during mobile-side deployment of AndesVL. Moreover, utilizing our cache eviction algorithm -- OKV -- along with customized speculative decoding and compression strategies, we achieve a 6.7x peak decoding speedup ratio, up to 30.9% memory reduction, and 1.8 bits-per-weight when deploying AndesVL-4B on MediaTek Dimensity 9500 chips. We release all models on https://huggingface.co/OPPOer.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11496
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model
Jin, Zhiwei
Song, Xiaohui
Wang, Nan
Liu, Yafei
Li, Chao
Li, Xin
Wang, Ruichen
Li, Zhihao
Qi, Qi
Cheng, Long
Hao, Dongze
Zheng, Quanlong
Zhang, Yanhao
Ji, Haobo
Ma, Jian
Zheng, Zhitong
Lin, Zhenyi
Deng, Haolin
Zou, Xin
Yin, Xiaojie
Wang, Ruilin
Cai, Liankai
Liu, Haijing
Qiu, Yuqing
Chen, Ke
Li, Zixian
Xie, Chi
Li, Huafei
Li, Chenxing
Wang, Chuangchuang
Tang, Kai
Zhu, Zhiguang
Tang, Kai
Gao, Wenmei
Wang, Rui
Wu, Jun
Liu, Chao
Xie, Qin
Chen, Chen
Lu, Haonan
Computer Vision and Pattern Recognition
Artificial Intelligence
In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hundreds of billions of parameters, they significantly surpass the limitations in memory, power consumption, and computing capacity of edge devices such as mobile phones. This paper introduces AndesVL, a suite of mobile-side MLLMs with 0.6B to 4B parameters based on Qwen3's LLM and various visual encoders. We comprehensively outline the model architectures, training pipeline, and training data of AndesVL, which achieves first-tier performance across a wide range of open-source benchmarks, including fields such as text-rich image understanding, reasoning and math, multi-image comprehension, general VQA, hallucination mitigation, multilingual understanding, and GUI-related tasks when compared with state-of-the-art models of a similar scale. Furthermore, we introduce a 1+N LoRA architecture alongside a Quantization-Aware LoRA Fine-Tuning (QALFT) framework to facilitate efficient task adaptation and model compression during mobile-side deployment of AndesVL. Moreover, utilizing our cache eviction algorithm -- OKV -- along with customized speculative decoding and compression strategies, we achieve a 6.7x peak decoding speedup ratio, up to 30.9% memory reduction, and 1.8 bits-per-weight when deploying AndesVL-4B on MediaTek Dimensity 9500 chips. We release all models on https://huggingface.co/OPPOer.
title AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.11496