Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ran, Dezhi, Wu, Mengzhou, Yu, Hao, Li, Yuetong, Ren, Jun, Cao, Yuan, Zeng, Xia, Lu, Haochuan, Xu, Zexin, Xu, Mengqian, Su, Ting, Yao, Liangchao, Xiong, Ting, Yang, Wei, Deng, Yuetang, Marron, Assaf, Harel, David, Xie, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Infrastructure Software Perspective Toward Computation Offloading between Executable Specifications and Foundation Models
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
A Specification's Realm: Characterizing the Knowledge Required for Executing a Given Algorithm Specification
von: Marron, Assaf, et al.
Veröffentlicht: (2025)
von: Marron, Assaf, et al.
Veröffentlicht: (2025)
Preparing for Super-Reactivity: Early Fault-Detection in the Development of Exceedingly Complex Reactive Systems
von: Harel, David, et al.
Veröffentlicht: (2024)
von: Harel, David, et al.
Veröffentlicht: (2024)
Improving Random Testing via LLM-powered UI Tarpit Escaping for Mobile Apps
von: Xu, Mengqian, et al.
Veröffentlicht: (2026)
von: Xu, Mengqian, et al.
Veröffentlicht: (2026)
Skill-Adpative Imitation Learning for UI Test Reuse
von: Wu, Mengzhou, et al.
Veröffentlicht: (2024)
von: Wu, Mengzhou, et al.
Veröffentlicht: (2024)
Enabling Cost-Effective UI Automation Testing with Retrieval-Based LLMs: A Case Study in WeChat
von: Feng, Sidong, et al.
Veröffentlicht: (2024)
von: Feng, Sidong, et al.
Veröffentlicht: (2024)
Foundation Model Engineering: Engineering Foundation Models Just as Engineering Software
von: Ran, Dezhi, et al.
Veröffentlicht: (2024)
von: Ran, Dezhi, et al.
Veröffentlicht: (2024)
UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics
von: Wu, Mengzhou, et al.
Veröffentlicht: (2026)
von: Wu, Mengzhou, et al.
Veröffentlicht: (2026)
Expecting the Unexpected: Developing Autonomous-System Design Principles for Reacting to Unpredicted Events and Conditions
von: Marron, Assaf, et al.
Veröffentlicht: (2020)
von: Marron, Assaf, et al.
Veröffentlicht: (2020)
Meta-autoencoders: An approach to discovery and representation of relationships between dynamically evolving classes
von: Marron, Assaf, et al.
Veröffentlicht: (2025)
von: Marron, Assaf, et al.
Veröffentlicht: (2025)
On Augmenting Scenario-Based Modeling with Generative AI
von: Harel, David, et al.
Veröffentlicht: (2024)
von: Harel, David, et al.
Veröffentlicht: (2024)
A Synthetic Pseudo-Autoencoder Invites Examination of Tacit Assumptions in Neural Network Design
von: Marron, Assaf
Veröffentlicht: (2025)
von: Marron, Assaf
Veröffentlicht: (2025)
Natural Averaging May Complement Known Biological Constraints in Bi-parental Reproduction's Advantages Over Mono-parental in Conserving Species Quantitative Traits
von: Marron, Assaf, et al.
Veröffentlicht: (2023)
von: Marron, Assaf, et al.
Veröffentlicht: (2023)
DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model
von: Yu, Hao, et al.
Veröffentlicht: (2025)
von: Yu, Hao, et al.
Veröffentlicht: (2025)
Evolution is Driven by Natural Autoencoding: Reframing Species, Interaction Codes, Cooperation, and Sexual Reproduction
von: Cohen, Irun R., et al.
Veröffentlicht: (2022)
von: Cohen, Irun R., et al.
Veröffentlicht: (2022)
Beyond Pass/Fail: The Story of Learning-Based Testing
von: Rahman, Sheikh Md. Mushfiqur, et al.
Veröffentlicht: (2025)
von: Rahman, Sheikh Md. Mushfiqur, et al.
Veröffentlicht: (2025)
iDiff: Interpretable Difference-aware Framework for Pairwise Image Quality Assessment
von: Yue, Xinli, et al.
Veröffentlicht: (2026)
von: Yue, Xinli, et al.
Veröffentlicht: (2026)
Data and System Perspectives of Sustainable Artificial Intelligence
von: Xie, Tao, et al.
Veröffentlicht: (2025)
von: Xie, Tao, et al.
Veröffentlicht: (2025)
GhostUI: Unveiling Hidden Interactions in Mobile UI
von: Kweon, Minkyu, et al.
Veröffentlicht: (2026)
von: Kweon, Minkyu, et al.
Veröffentlicht: (2026)
MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding
von: Parvez, Athar, et al.
Veröffentlicht: (2026)
von: Parvez, Athar, et al.
Veröffentlicht: (2026)
Identifying User Goals from UI Trajectories
von: Berkovitch, Omri, et al.
Veröffentlicht: (2024)
von: Berkovitch, Omri, et al.
Veröffentlicht: (2024)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
von: You, Keen, et al.
Veröffentlicht: (2024)
von: You, Keen, et al.
Veröffentlicht: (2024)
JSidentify-V2: Leveraging Dynamic Memory Fingerprinting for Mini-Game Plagiarism Detection
von: Li, Zhihao, et al.
Veröffentlicht: (2025)
von: Li, Zhihao, et al.
Veröffentlicht: (2025)
MANA: Towards Efficient Mobile Ad Detection via Multimodal Agentic UI Navigation
von: Zhao, Yizhe, et al.
Veröffentlicht: (2026)
von: Zhao, Yizhe, et al.
Veröffentlicht: (2026)
Scalable Density-based Clustering with Random Projections
von: Xu, Haochuan, et al.
Veröffentlicht: (2024)
von: Xu, Haochuan, et al.
Veröffentlicht: (2024)
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2025)
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2025)
Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
von: Zhou, Zirun, et al.
Veröffentlicht: (2025)
von: Zhou, Zirun, et al.
Veröffentlicht: (2025)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
von: Xu, Qiao, et al.
Veröffentlicht: (2026)
von: Xu, Qiao, et al.
Veröffentlicht: (2026)
UniGoal: Towards Universal Zero-shot Goal-oriented Navigation
von: Yin, Hang, et al.
Veröffentlicht: (2025)
von: Yin, Hang, et al.
Veröffentlicht: (2025)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
von: Lin, Zichuan, et al.
Veröffentlicht: (2026)
von: Lin, Zichuan, et al.
Veröffentlicht: (2026)
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
von: Zhang, Li, et al.
Veröffentlicht: (2024)
von: Zhang, Li, et al.
Veröffentlicht: (2024)
Trade‐Offs Between Economic Gains and Ecological Goals: Impact of Marine Ecological Policy on Marine Fisheries Efficiency
von: Jintao Ma, et al.
Veröffentlicht: (2025)
von: Jintao Ma, et al.
Veröffentlicht: (2025)
An Empirical Study and Theoretical Explanation on Task-Level Model-Merging Collapse
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
Message Passing Based Demodulation of the Time-Encoded Digital Modulation Signal
von: Xu, Yuan, et al.
Veröffentlicht: (2025)
von: Xu, Yuan, et al.
Veröffentlicht: (2025)
A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation Models
von: Yu, Rosen Ting-Ying, et al.
Veröffentlicht: (2025)
von: Yu, Rosen Ting-Ying, et al.
Veröffentlicht: (2025)
Beyond Isolation: A Unified Benchmark for General-Purpose Navigation
von: Sun, Samson, et al.
Veröffentlicht: (2026)
von: Sun, Samson, et al.
Veröffentlicht: (2026)
User-Centric Design of UI for Mobile Banking Apps: Improving UI and Features for Better Customer Experience
von: Chitrakar, Luniva, et al.
Veröffentlicht: (2026)
von: Chitrakar, Luniva, et al.
Veröffentlicht: (2026)
UI Semantic Group Detection: Grouping UI Elements with Similar Semantics in Mobile Graphical User Interface
von: Xiao, Shuhong, et al.
Veröffentlicht: (2024)
von: Xiao, Shuhong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An Infrastructure Software Perspective Toward Computation Offloading between Executable Specifications and Foundation Models
von: Ran, Dezhi, et al.
Veröffentlicht: (2025) -
From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
von: Ran, Dezhi, et al.
Veröffentlicht: (2025) -
A Specification's Realm: Characterizing the Knowledge Required for Executing a Given Algorithm Specification
von: Marron, Assaf, et al.
Veröffentlicht: (2025) -
Preparing for Super-Reactivity: Early Fault-Detection in the Development of Exceedingly Complex Reactive Systems
von: Harel, David, et al.
Veröffentlicht: (2024) -
Improving Random Testing via LLM-powered UI Tarpit Escaping for Mobile Apps
von: Xu, Mengqian, et al.
Veröffentlicht: (2026)