Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Jihao, Ai, Qihang, Wang, Yingyao, Bu, Pi, Xing, Jingxuan, Zhu, Zekun, Jiang, Wei, Wang, Ziming, Zhao, Yingxiu, Zhang, Ming-Liang, Song, Jun, Jiang, Yuning, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
by: Ai, Qihang, et al.
Published: (2025)
by: Ai, Qihang, et al.
Published: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
by: Xie, Yiping, et al.
Published: (2026)
by: Xie, Yiping, et al.
Published: (2026)
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
by: Cao, Yue, et al.
Published: (2025)
by: Cao, Yue, et al.
Published: (2025)
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
by: Gu, Jihao, et al.
Published: (2024)
by: Gu, Jihao, et al.
Published: (2024)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
Systematic investigation of trace anomaly contribution in nucleon mass
by: Wang, Xiao-Yun, et al.
Published: (2024)
by: Wang, Xiao-Yun, et al.
Published: (2024)
Generalization in Online Reinforcement Learning for Mobile Agents
by: Gu, Li, et al.
Published: (2026)
by: Gu, Li, et al.
Published: (2026)
Parallel Thinking, Sequential Answering: Bridging NAR and AR for Efficient Reasoning
by: Ai, Qihang, et al.
Published: (2025)
by: Ai, Qihang, et al.
Published: (2025)
PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
by: Gong, Yichen, et al.
Published: (2026)
by: Gong, Yichen, et al.
Published: (2026)
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
by: Chu, Xiangxiang, et al.
Published: (2023)
by: Chu, Xiangxiang, et al.
Published: (2023)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
AppAgent v2: Advanced Agent for Flexible Mobile Interactions
by: Li, Yanda, et al.
Published: (2024)
by: Li, Yanda, et al.
Published: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training
by: Kang, Linjia, et al.
Published: (2026)
by: Kang, Linjia, et al.
Published: (2026)
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
by: Kong, Quyu, et al.
Published: (2025)
by: Kong, Quyu, et al.
Published: (2025)
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
by: Wang, Xuan, et al.
Published: (2025)
by: Wang, Xuan, et al.
Published: (2025)
CRUD-Capable Mobile Apps with R and shinyMobile: a Case Study in Rapid Prototyping
by: Henry, Nathan
Published: (2024)
by: Henry, Nathan
Published: (2024)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
by: Wang, Zekun, et al.
Published: (2025)
by: Wang, Zekun, et al.
Published: (2025)
How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Graph-to-Vision: Multi-graph Understanding and Reasoning using Vision-Language Models
by: Ai, Qihang, et al.
Published: (2025)
by: Ai, Qihang, et al.
Published: (2025)
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
by: Huang, Kun, et al.
Published: (2025)
by: Huang, Kun, et al.
Published: (2025)
Embodied AI-Enhanced IoMT Edge Computing: UAV Trajectory Optimization and Task Offloading with Mobility Prediction
by: Mu, Siqi, et al.
Published: (2025)
by: Mu, Siqi, et al.
Published: (2025)
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
by: Wang, Junyang, et al.
Published: (2026)
by: Wang, Junyang, et al.
Published: (2026)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
by: Wang, Luyuan, et al.
Published: (2024)
by: Wang, Luyuan, et al.
Published: (2024)
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
MobiAgent: A Systematic Framework for Customizable Mobile Agents
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
Mobile Cell-Free Massive MIMO with Multi-Agent Reinforcement Learning: A Scalable Framework
by: Liu, Ziheng, et al.
Published: (2024)
by: Liu, Ziheng, et al.
Published: (2024)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence
by: He, Bowei
Published: (2026)
by: He, Bowei
Published: (2026)
ReStory: VLM-augmentation of Social Human-Robot Interaction Datasets
by: Bu, Fanjun, et al.
Published: (2024)
by: Bu, Fanjun, et al.
Published: (2024)
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
ICPRL: Acquiring Physical Intuition from Interactive Control
by: Xu, Xinrun, et al.
Published: (2026)
by: Xu, Xinrun, et al.
Published: (2026)
MobileExperts: A Dynamic Tool-Enabled Agent Team in Mobile Devices
by: Zhang, Jiayi, et al.
Published: (2024)
by: Zhang, Jiayi, et al.
Published: (2024)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
by: Wang, Junyang, et al.
Published: (2025)
by: Wang, Junyang, et al.
Published: (2025)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
by: Wang, Junyang, et al.
Published: (2025)
by: Wang, Junyang, et al.
Published: (2025)
Similar Items
-
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
by: Ai, Qihang, et al.
Published: (2025) -
SecAgent: Efficient Mobile GUI Agent with Semantic Context
by: Xie, Yiping, et al.
Published: (2026) -
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
by: Cao, Yue, et al.
Published: (2025) -
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
by: Xu, Liangyu, et al.
Published: (2025) -
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
by: Gu, Jihao, et al.
Published: (2024)