SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866918330834616320 |
|---|---|
| author | Tian, Muxin Wang, Zhe Yang, Blair Tang, Zhenwei Zhu, Kunlun Dong, Honghua Li, Hanchen Xie, Xinni Wang, Guangjing You, Jiaxuan |
| author_facet | Tian, Muxin Wang, Zhe Yang, Blair Tang, Zhenwei Zhu, Kunlun Dong, Honghua Li, Hanchen Xie, Xinni Wang, Guangjing You, Jiaxuan |
| contents | Can large language model agents develop industry-level mobile applications? We introduce \textbf{SWE-Bench Mobile}, a benchmark for evaluating coding agents on realistic software engineering tasks derived from a production iOS codebase. Unlike existing benchmarks that focus on isolated problems or bug fixes, SWE-Bench Mobile captures the full complexity of industrial development: multi-modal inputs (PRDs and Figma designs), a large-scale mixed Swift/Objective-C codebase, and comprehensive test suites. We evaluate 22 agent-model configurations across four coding agents -- three commercial (Cursor, Codex, Claude Code) and one open-source (OpenCode) -- and find that even the best configurations achieve only 12\% task success rate. Our analysis reveals that (1) agent design matters as much as model capability -- the same model shows up to 6$\times$ performance gap across agents, (2) commercial agents consistently outperform open-source alternatives, and (3) simple ``Defensive Programming'' prompts outperform complex ones by 7.4\%. These findings highlight a significant gap between current agent capabilities and industrial requirements, while providing actionable insights for practitioners and researchers. We release SWE-Bench Mobile as a \textit{hosted benchmark challenge} to prevent data contamination and ensure fair evaluation. The public leaderboard and development toolkit are available at https://swebenchmobile.com. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_09540 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications? Tian, Muxin Wang, Zhe Yang, Blair Tang, Zhenwei Zhu, Kunlun Dong, Honghua Li, Hanchen Xie, Xinni Wang, Guangjing You, Jiaxuan Software Engineering Can large language model agents develop industry-level mobile applications? We introduce \textbf{SWE-Bench Mobile}, a benchmark for evaluating coding agents on realistic software engineering tasks derived from a production iOS codebase. Unlike existing benchmarks that focus on isolated problems or bug fixes, SWE-Bench Mobile captures the full complexity of industrial development: multi-modal inputs (PRDs and Figma designs), a large-scale mixed Swift/Objective-C codebase, and comprehensive test suites. We evaluate 22 agent-model configurations across four coding agents -- three commercial (Cursor, Codex, Claude Code) and one open-source (OpenCode) -- and find that even the best configurations achieve only 12\% task success rate. Our analysis reveals that (1) agent design matters as much as model capability -- the same model shows up to 6$\times$ performance gap across agents, (2) commercial agents consistently outperform open-source alternatives, and (3) simple ``Defensive Programming'' prompts outperform complex ones by 7.4\%. These findings highlight a significant gap between current agent capabilities and industrial requirements, while providing actionable insights for practitioners and researchers. We release SWE-Bench Mobile as a \textit{hosted benchmark challenge} to prevent data contamination and ensure fair evaluation. The public leaderboard and development toolkit are available at https://swebenchmobile.com. |
| title | SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications? |
| topic | Software Engineering |
| url | https://arxiv.org/abs/2602.09540 |