SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tian, Muxin, Wang, Zhe, Yang, Blair, Tang, Zhenwei, Zhu, Kunlun, Dong, Honghua, Li, Hanchen, Xie, Xinni, Wang, Guangjing, You, Jiaxuan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918330834616320
author Tian, Muxin
Wang, Zhe
Yang, Blair
Tang, Zhenwei
Zhu, Kunlun
Dong, Honghua
Li, Hanchen
Xie, Xinni
Wang, Guangjing
You, Jiaxuan
author_facet Tian, Muxin
Wang, Zhe
Yang, Blair
Tang, Zhenwei
Zhu, Kunlun
Dong, Honghua
Li, Hanchen
Xie, Xinni
Wang, Guangjing
You, Jiaxuan
contents Can large language model agents develop industry-level mobile applications? We introduce \textbf{SWE-Bench Mobile}, a benchmark for evaluating coding agents on realistic software engineering tasks derived from a production iOS codebase. Unlike existing benchmarks that focus on isolated problems or bug fixes, SWE-Bench Mobile captures the full complexity of industrial development: multi-modal inputs (PRDs and Figma designs), a large-scale mixed Swift/Objective-C codebase, and comprehensive test suites. We evaluate 22 agent-model configurations across four coding agents -- three commercial (Cursor, Codex, Claude Code) and one open-source (OpenCode) -- and find that even the best configurations achieve only 12\% task success rate. Our analysis reveals that (1) agent design matters as much as model capability -- the same model shows up to 6$\times$ performance gap across agents, (2) commercial agents consistently outperform open-source alternatives, and (3) simple ``Defensive Programming'' prompts outperform complex ones by 7.4\%. These findings highlight a significant gap between current agent capabilities and industrial requirements, while providing actionable insights for practitioners and researchers. We release SWE-Bench Mobile as a \textit{hosted benchmark challenge} to prevent data contamination and ensure fair evaluation. The public leaderboard and development toolkit are available at https://swebenchmobile.com.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09540
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?
Tian, Muxin
Wang, Zhe
Yang, Blair
Tang, Zhenwei
Zhu, Kunlun
Dong, Honghua
Li, Hanchen
Xie, Xinni
Wang, Guangjing
You, Jiaxuan
Software Engineering
Can large language model agents develop industry-level mobile applications? We introduce \textbf{SWE-Bench Mobile}, a benchmark for evaluating coding agents on realistic software engineering tasks derived from a production iOS codebase. Unlike existing benchmarks that focus on isolated problems or bug fixes, SWE-Bench Mobile captures the full complexity of industrial development: multi-modal inputs (PRDs and Figma designs), a large-scale mixed Swift/Objective-C codebase, and comprehensive test suites. We evaluate 22 agent-model configurations across four coding agents -- three commercial (Cursor, Codex, Claude Code) and one open-source (OpenCode) -- and find that even the best configurations achieve only 12\% task success rate. Our analysis reveals that (1) agent design matters as much as model capability -- the same model shows up to 6$\times$ performance gap across agents, (2) commercial agents consistently outperform open-source alternatives, and (3) simple ``Defensive Programming'' prompts outperform complex ones by 7.4\%. These findings highlight a significant gap between current agent capabilities and industrial requirements, while providing actionable insights for practitioners and researchers. We release SWE-Bench Mobile as a \textit{hosted benchmark challenge} to prevent data contamination and ensure fair evaluation. The public leaderboard and development toolkit are available at https://swebenchmobile.com.
title SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?
topic Software Engineering
url https://arxiv.org/abs/2602.09540