Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Junyang, Xu, Haiyang, Zhang, Xi, Yan, Ming, Zhang, Ji, Huang, Fei, Sang, Jitao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
von: Wang, Junyang, et al.
Veröffentlicht: (2025)
von: Wang, Junyang, et al.
Veröffentlicht: (2025)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
von: Wang, Junyang, et al.
Veröffentlicht: (2026)
von: Wang, Junyang, et al.
Veröffentlicht: (2026)
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
von: Liu, Haowei, et al.
Veröffentlicht: (2025)
von: Liu, Haowei, et al.
Veröffentlicht: (2025)
FairCLIP: Social Bias Elimination based on Attribute Prototype Learning and Representation Neutralization
von: Wang, Junyang, et al.
Veröffentlicht: (2022)
von: Wang, Junyang, et al.
Veröffentlicht: (2022)
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
von: Xu, Haiyang, et al.
Veröffentlicht: (2026)
von: Xu, Haiyang, et al.
Veröffentlicht: (2026)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
von: Wang, Junyang, et al.
Veröffentlicht: (2023)
von: Wang, Junyang, et al.
Veröffentlicht: (2023)
Mobile-Agent-v3: Fundamental Agents for GUI Automation
von: Ye, Jiabo, et al.
Veröffentlicht: (2025)
von: Ye, Jiabo, et al.
Veröffentlicht: (2025)
Proxy Robustness in Vision Language Models is Effortlessly Transferable
von: Fu, Xiaowei, et al.
Veröffentlicht: (2026)
von: Fu, Xiaowei, et al.
Veröffentlicht: (2026)
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
von: Zhang, Liang, et al.
Veröffentlicht: (2024)
von: Zhang, Liang, et al.
Veröffentlicht: (2024)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2026)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2026)
Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
von: Chen, Gehui, et al.
Veröffentlicht: (2025)
von: Chen, Gehui, et al.
Veröffentlicht: (2025)
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
von: Zhang, Shuai, et al.
Veröffentlicht: (2025)
von: Zhang, Shuai, et al.
Veröffentlicht: (2025)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
von: Loo, Gowen, et al.
Veröffentlicht: (2025)
von: Loo, Gowen, et al.
Veröffentlicht: (2025)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
von: Guo, Ruoqi, et al.
Veröffentlicht: (2026)
von: Guo, Ruoqi, et al.
Veröffentlicht: (2026)
MobileExperts: A Dynamic Tool-Enabled Agent Team in Mobile Devices
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
von: Kong, Quyu, et al.
Veröffentlicht: (2025)
von: Kong, Quyu, et al.
Veröffentlicht: (2025)
MobileFlow: A Multimodal LLM For Mobile GUI Agent
von: Nong, Songqin, et al.
Veröffentlicht: (2024)
von: Nong, Songqin, et al.
Veröffentlicht: (2024)
MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
von: Kong, Yi, et al.
Veröffentlicht: (2025)
von: Kong, Yi, et al.
Veröffentlicht: (2025)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
von: Yang, Min, et al.
Veröffentlicht: (2025)
von: Yang, Min, et al.
Veröffentlicht: (2025)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
Trinational Automated Mobility
von: Vogt, Jonas, et al.
Veröffentlicht: (2021)
von: Vogt, Jonas, et al.
Veröffentlicht: (2021)
MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
MobiAct: Efficient MAV Action Recognition Using MobileNetV4 with Contrastive Learning and Knowledge Distillation
von: Nengbo, Zhang, et al.
Veröffentlicht: (2025)
von: Nengbo, Zhang, et al.
Veröffentlicht: (2025)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2025)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2025)
From Human Negotiation to Agent Negotiation: Personal Mobility Agents in Automated Traffic
von: Jansen, Pascal
Veröffentlicht: (2026)
von: Jansen, Pascal
Veröffentlicht: (2026)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
von: Li, Senyao, et al.
Veröffentlicht: (2026)
von: Li, Senyao, et al.
Veröffentlicht: (2026)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
von: Jang, Yunseok, et al.
Veröffentlicht: (2025)
von: Jang, Yunseok, et al.
Veröffentlicht: (2025)
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
von: Huang, Kun, et al.
Veröffentlicht: (2025)
von: Huang, Kun, et al.
Veröffentlicht: (2025)
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
von: Guo, Yuan, et al.
Veröffentlicht: (2025)
von: Guo, Yuan, et al.
Veröffentlicht: (2025)
MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
Flora: Effortless Context Construction to Arbitrary Length and Scale
von: Chen, Tianxiang, et al.
Veröffentlicht: (2025)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
von: Wang, Junyang, et al.
Veröffentlicht: (2025) -
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
von: Wang, Junyang, et al.
Veröffentlicht: (2024) -
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024) -
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025) -
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
von: Wang, Junyang, et al.
Veröffentlicht: (2026)