MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Qinzhuo, Xu, Weikai, Liu, Wei, Tan, Tao, Liu, Jianfeng, Li, Ang, Luan, Jian, Wang, Bin, Shang, Shuo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
por: Chu, Xiangxiang, et al.
Publicado: (2023)
por: Chu, Xiangxiang, et al.
Publicado: (2023)
ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation
por: Wu, Qinzhuo, et al.
Publicado: (2025)
por: Wu, Qinzhuo, et al.
Publicado: (2025)
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
por: Chu, Xiangxiang, et al.
Publicado: (2024)
por: Chu, Xiangxiang, et al.
Publicado: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
por: Deng, Shihan, et al.
Publicado: (2024)
por: Deng, Shihan, et al.
Publicado: (2024)
ToolPlanner: A Tool Augmented LLM for Multi Granularity Instructions with Path Planning and Feedback
por: Wu, Qinzhuo, et al.
Publicado: (2024)
por: Wu, Qinzhuo, et al.
Publicado: (2024)
Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
por: Xu, Weikai, et al.
Publicado: (2025)
por: Xu, Weikai, et al.
Publicado: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
por: Wu, Qinzhuo, et al.
Publicado: (2026)
por: Wu, Qinzhuo, et al.
Publicado: (2026)
BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
por: Wu, Qinzhuo, et al.
Publicado: (2025)
por: Wu, Qinzhuo, et al.
Publicado: (2025)
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
por: Huang, Kun, et al.
Publicado: (2025)
por: Huang, Kun, et al.
Publicado: (2025)
DetermLR: Augmenting LLM-based Logical Reasoning from Indeterminacy to Determinacy
por: Sun, Hongda, et al.
Publicado: (2023)
por: Sun, Hongda, et al.
Publicado: (2023)
R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
por: Jiang, Zhizheng, et al.
Publicado: (2026)
por: Jiang, Zhizheng, et al.
Publicado: (2026)
CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning
por: Liu, Yuxuan, et al.
Publicado: (2026)
por: Liu, Yuxuan, et al.
Publicado: (2026)
PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning
por: Wang, Qibin, et al.
Publicado: (2024)
por: Wang, Qibin, et al.
Publicado: (2024)
Improving Random Testing via LLM-powered UI Tarpit Escaping for Mobile Apps
por: Xu, Mengqian, et al.
Publicado: (2026)
por: Xu, Mengqian, et al.
Publicado: (2026)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
por: You, Keen, et al.
Publicado: (2024)
por: You, Keen, et al.
Publicado: (2024)
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
por: Qu, Heng, et al.
Publicado: (2026)
por: Qu, Heng, et al.
Publicado: (2026)
Pruning Large Language Models to Intra-module Low-rank Architecture with Transitional Activations
por: Shen, Bowen, et al.
Publicado: (2024)
por: Shen, Bowen, et al.
Publicado: (2024)
User-Centric Design of UI for Mobile Banking Apps: Improving UI and Features for Better Customer Experience
por: Chitrakar, Luniva, et al.
Publicado: (2026)
por: Chitrakar, Luniva, et al.
Publicado: (2026)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
por: Liu, Hanqing, et al.
Publicado: (2026)
por: Liu, Hanqing, et al.
Publicado: (2026)
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
por: Chen, Yuhan, et al.
Publicado: (2024)
por: Chen, Yuhan, et al.
Publicado: (2024)
VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
por: Song, Yunpeng, et al.
Publicado: (2023)
por: Song, Yunpeng, et al.
Publicado: (2023)
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
por: Wang, Dianyi, et al.
Publicado: (2025)
por: Wang, Dianyi, et al.
Publicado: (2025)
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
por: Baechler, Gilles, et al.
Publicado: (2024)
por: Baechler, Gilles, et al.
Publicado: (2024)
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
por: Gao, Longxi, et al.
Publicado: (2025)
por: Gao, Longxi, et al.
Publicado: (2025)
Modeling Inter-Intra Heterogeneity for Graph Federated Learning
por: Yu, Wentao, et al.
Publicado: (2024)
por: Yu, Wentao, et al.
Publicado: (2024)
More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
por: Zhang, Xiaoqing, et al.
Publicado: (2025)
por: Zhang, Xiaoqing, et al.
Publicado: (2025)
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
por: Tian, Beitong, et al.
Publicado: (2025)
por: Tian, Beitong, et al.
Publicado: (2025)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
por: Liu, Jiajin, et al.
Publicado: (2026)
por: Liu, Jiajin, et al.
Publicado: (2026)
From UI to Code: Mobile Ads Detection via LLM-Unified Static-Dynamic Analysis
por: Ma, Shang, et al.
Publicado: (2026)
por: Ma, Shang, et al.
Publicado: (2026)
InterVLS: Interactive Model Understanding and Improvement with Vision-Language Surrogates
por: Huang, Jinbin, et al.
Publicado: (2023)
por: Huang, Jinbin, et al.
Publicado: (2023)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
por: Masry, Ahmed, et al.
Publicado: (2025)
por: Masry, Ahmed, et al.
Publicado: (2025)
OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference
por: Chen, Wei, et al.
Publicado: (2024)
por: Chen, Wei, et al.
Publicado: (2024)
Theoretical Investigations of Intra‐ and Inter‐ Interactions of Wogonin and Wogonoside
por: Xiong Li, et al.
Publicado: (2025)
por: Xiong Li, et al.
Publicado: (2025)
UI-UG: A Unified MLLM for UI Understanding and Generation
por: Yang, Hao, et al.
Publicado: (2025)
por: Yang, Hao, et al.
Publicado: (2025)
BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning
por: Ke, Jingyang, et al.
Publicado: (2026)
por: Ke, Jingyang, et al.
Publicado: (2026)
GhostUI: Unveiling Hidden Interactions in Mobile UI
por: Kweon, Minkyu, et al.
Publicado: (2026)
por: Kweon, Minkyu, et al.
Publicado: (2026)
VLM-Fuzz: Vision Language Model Assisted Recursive Depth-first Search Exploration for Effective UI Testing of Android Apps
por: Demissie, Biniam Fisseha, et al.
Publicado: (2025)
por: Demissie, Biniam Fisseha, et al.
Publicado: (2025)
Scaling Model and Data for Multilingual Machine Translation with Open Large Language Models
por: Shang, Yuzhe, et al.
Publicado: (2026)
por: Shang, Yuzhe, et al.
Publicado: (2026)
MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding
por: Parvez, Athar, et al.
Publicado: (2026)
por: Parvez, Athar, et al.
Publicado: (2026)
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
por: Chen, Pingyi, et al.
Publicado: (2025)
por: Chen, Pingyi, et al.
Publicado: (2025)
Ejemplares similares
-
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
por: Chu, Xiangxiang, et al.
Publicado: (2023) -
ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation
por: Wu, Qinzhuo, et al.
Publicado: (2025) -
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
por: Chu, Xiangxiang, et al.
Publicado: (2024) -
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
por: Deng, Shihan, et al.
Publicado: (2024) -
ToolPlanner: A Tool Augmented LLM for Multi Granularity Instructions with Path Planning and Feedback
por: Wu, Qinzhuo, et al.
Publicado: (2024)