MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Xiangxiang, Qiao, Limeng, Lin, Xinyang, Xu, Shuang, Yang, Yang, Hu, Yiming, Wei, Fei, Zhang, Xinyu, Zhang, Bo, Wei, Xiaolin, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
by: Wu, Qinzhuo, et al.
Published: (2024)
by: Wu, Qinzhuo, et al.
Published: (2024)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
by: Liu, Zhaochen, et al.
Published: (2024)
by: Liu, Zhaochen, et al.
Published: (2024)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
by: Chu, Xiangxiang, et al.
Published: (2025)
by: Chu, Xiangxiang, et al.
Published: (2025)
Agent-SAMA: State-Aware Mobile Assistant
by: Guo, Linqiang, et al.
Published: (2025)
by: Guo, Linqiang, et al.
Published: (2025)
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices
by: Wang, Zhaode, et al.
Published: (2025)
by: Wang, Zhaode, et al.
Published: (2025)
VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
by: Chen, Zizhi, et al.
Published: (2025)
by: Chen, Zizhi, et al.
Published: (2025)
PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
by: Lu, Xudong, et al.
Published: (2026)
by: Lu, Xudong, et al.
Published: (2026)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices
by: Yan, Yongsheng, et al.
Published: (2026)
by: Yan, Yongsheng, et al.
Published: (2026)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
by: Xue, Xizhe, et al.
Published: (2024)
by: Xue, Xizhe, et al.
Published: (2024)
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
by: Zhang, Shuai, et al.
Published: (2025)
by: Zhang, Shuai, et al.
Published: (2025)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
FastGrasp: Learning-based Whole-body Control method for Fast Dexterous Grasping with Mobile Manipulators
by: Tao, Heng, et al.
Published: (2026)
by: Tao, Heng, et al.
Published: (2026)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
by: Song, Zhiheng, et al.
Published: (2026)
by: Song, Zhiheng, et al.
Published: (2026)
UniViTAR: Unified Vision Transformer with Native Resolution
by: Qiao, Limeng, et al.
Published: (2025)
by: Qiao, Limeng, et al.
Published: (2025)
MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
by: Zhao, Yang, et al.
Published: (2023)
by: Zhao, Yang, et al.
Published: (2023)
Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
by: Wu, Zhe, et al.
Published: (2025)
by: Wu, Zhe, et al.
Published: (2025)
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
by: Zou, Ya, et al.
Published: (2025)
by: Zou, Ya, et al.
Published: (2025)
BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models
by: Deng, Fei, et al.
Published: (2026)
by: Deng, Fei, et al.
Published: (2026)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
by: Yang, Min, et al.
Published: (2025)
by: Yang, Min, et al.
Published: (2025)
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
by: Ai, Qihang, et al.
Published: (2025)
by: Ai, Qihang, et al.
Published: (2025)
Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
Defending against Patch-Based and Texture-Based Adversarial Attacks with Spectral Decomposition
by: Zhang, Wei, et al.
Published: (2026)
by: Zhang, Wei, et al.
Published: (2026)
Real-Time Vehicle Detection and Urban Traffic Behavior Analysis Based on UAV Traffic Videos on Mobile Devices
by: Zhu, Yuan, et al.
Published: (2024)
by: Zhu, Yuan, et al.
Published: (2024)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
by: Li, YuQian, et al.
Published: (2025)
by: Li, YuQian, et al.
Published: (2025)
Memory-Efficient Backpropagation for Fine-Tuning LLMs on Resource-Constrained Mobile Devices
by: Song, Congzheng, et al.
Published: (2025)
by: Song, Congzheng, et al.
Published: (2025)
ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices
by: Kong, Dezhi, et al.
Published: (2026)
by: Kong, Dezhi, et al.
Published: (2026)
Diatom‐Inspired 1D Immobile Robots Capable of 2D Collective Mobility
by: Tianyi Hu, et al.
Published: (2026)
by: Tianyi Hu, et al.
Published: (2026)
Adaptive Offloading and Enhancement for Low-Light Video Analytics on Mobile Devices
by: He, Yuanyi, et al.
Published: (2024)
by: He, Yuanyi, et al.
Published: (2024)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
by: Cheng, Kanzhi, et al.
Published: (2026)
by: Cheng, Kanzhi, et al.
Published: (2026)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
A High‐Mobility n‐Type Noncovalently‐Fused‐Ring Polymer for High‐Performance Organic Thermoelectrics
by: Tao Shen, et al.
Published: (2024)
by: Tao Shen, et al.
Published: (2024)
MobileExperts: A Dynamic Tool-Enabled Agent Team in Mobile Devices
by: Zhang, Jiayi, et al.
Published: (2024)
by: Zhang, Jiayi, et al.
Published: (2024)
Similar Items
-
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
by: Chu, Xiangxiang, et al.
Published: (2024) -
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
by: Wu, Qinzhuo, et al.
Published: (2024) -
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
by: Chu, Xiangxiang, et al.
Published: (2024) -
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
by: Wang, Junyang, et al.
Published: (2024) -
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
by: Liu, Zhaochen, et al.
Published: (2024)