MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Yunfei, Zhao, Xi, Zhang, Cheng, Feng, Dahu, Cheng, Daolin, Yu, Jianqi, Xia, Yubin, Feng, Erhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SkVM: Revisiting Language VM for Skills across Heterogenous LLMs and Harnesses
von: Chen, Le, et al.
Veröffentlicht: (2026)
von: Chen, Le, et al.
Veröffentlicht: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
von: Xu, Qiao, et al.
Veröffentlicht: (2026)
von: Xu, Qiao, et al.
Veröffentlicht: (2026)
MobiAgent: A Systematic Framework for Customizable Mobile Agents
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
von: Zhang, Cheng, et al.
Veröffentlicht: (2025)
Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
von: Liu, Zibin, et al.
Veröffentlicht: (2025)
von: Liu, Zibin, et al.
Veröffentlicht: (2025)
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
von: Hu, Li, et al.
Veröffentlicht: (2025)
von: Hu, Li, et al.
Veröffentlicht: (2025)
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
von: Zhu, Tianhao, et al.
Veröffentlicht: (2025)
von: Zhu, Tianhao, et al.
Veröffentlicht: (2025)
PanicFI: An Infrastructure for Fixing Panic Bugs in Real-World Rust Programs
von: Ni, Yunbo, et al.
Veröffentlicht: (2024)
von: Ni, Yunbo, et al.
Veröffentlicht: (2024)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
A Benchmark for Language Models in Real-World System Building
von: Jin, Weilin, et al.
Veröffentlicht: (2026)
von: Jin, Weilin, et al.
Veröffentlicht: (2026)
PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
von: Wei, Zichao, et al.
Veröffentlicht: (2025)
von: Wei, Zichao, et al.
Veröffentlicht: (2025)
MobileUPReg: Identifying User-Perceived Performance Regressions in Mobile OS Versions
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
von: Leung, Ho Fai, et al.
Veröffentlicht: (2025)
von: Leung, Ho Fai, et al.
Veröffentlicht: (2025)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
von: Feng, Jia, et al.
Veröffentlicht: (2024)
von: Feng, Jia, et al.
Veröffentlicht: (2024)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
GUIWatcher: Automatically Detecting GUI Lags by Analyzing Mobile Application Screencasts
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
von: Liu, Shuhan, et al.
Veröffentlicht: (2026)
von: Liu, Shuhan, et al.
Veröffentlicht: (2026)
FuncDroid: Towards Inter-Functional Flows for Comprehensive Mobile App GUI Testing
von: He, Jinlong, et al.
Veröffentlicht: (2026)
von: He, Jinlong, et al.
Veröffentlicht: (2026)
MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios
von: Huang, Jinyang, et al.
Veröffentlicht: (2025)
von: Huang, Jinyang, et al.
Veröffentlicht: (2025)
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
von: Ahmed, Md Basim Uddin, et al.
Veröffentlicht: (2025)
von: Ahmed, Md Basim Uddin, et al.
Veröffentlicht: (2025)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
von: Ahmed, Syed Yusuf, et al.
Veröffentlicht: (2026)
von: Ahmed, Syed Yusuf, et al.
Veröffentlicht: (2026)
MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
APPATCH: Automated Adaptive Prompting Large Language Models for Real-World Software Vulnerability Patching
von: Nong, Yu, et al.
Veröffentlicht: (2024)
von: Nong, Yu, et al.
Veröffentlicht: (2024)
SEAlign: Alignment Training for Software Engineering Agent
von: Zhang, Kechi, et al.
Veröffentlicht: (2025)
von: Zhang, Kechi, et al.
Veröffentlicht: (2025)
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
von: Jain, Kush, et al.
Veröffentlicht: (2024)
von: Jain, Kush, et al.
Veröffentlicht: (2024)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
MultiTest: Physical-Aware Object Insertion for Testing Multi-sensor Fusion Perception Systems
von: Gao, Xinyu, et al.
Veröffentlicht: (2024)
von: Gao, Xinyu, et al.
Veröffentlicht: (2024)
Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
LLM-CompDroid: Repairing Configuration Compatibility Bugs in Android Apps with Pre-trained Large Language Models
von: Liu, Zhijie, et al.
Veröffentlicht: (2024)
von: Liu, Zhijie, et al.
Veröffentlicht: (2024)
Dependency-Guided Repository-Level C-to-Rust Translation with Reinforcement Alignment
von: Feng, Jia, et al.
Veröffentlicht: (2026)
von: Feng, Jia, et al.
Veröffentlicht: (2026)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
von: Zhang, Kechi, et al.
Veröffentlicht: (2024)
von: Zhang, Kechi, et al.
Veröffentlicht: (2024)
A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
SWE-Fuse: Empowering Software Agents via Issue-free Trajectory Learning and Entropy-aware RLVR Training
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2026)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2026)
Integrating Rules and Semantics for LLM-Based C-to-Rust Translation
von: Luo, Feng, et al.
Veröffentlicht: (2025)
von: Luo, Feng, et al.
Veröffentlicht: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
Don't Confuse! Redrawing GUI Navigation Flow in Mobile Apps for Visually Impaired Users
von: Zhang, Mengxi, et al.
Veröffentlicht: (2025)
von: Zhang, Mengxi, et al.
Veröffentlicht: (2025)
Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SkVM: Revisiting Language VM for Skills across Heterogenous LLMs and Harnesses
von: Chen, Le, et al.
Veröffentlicht: (2026) -
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
von: Xu, Qiao, et al.
Veröffentlicht: (2026) -
MobiAgent: A Systematic Framework for Customizable Mobile Agents
von: Zhang, Cheng, et al.
Veröffentlicht: (2025) -
Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
von: Liu, Zibin, et al.
Veröffentlicht: (2025) -
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
von: Hu, Li, et al.
Veröffentlicht: (2025)