GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xiaoyi, Gao, Yifei, Xu, Yang, Song, Xingxing, Zhang, Yi, Sang, Jitao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
von: Bian, Yutong, et al.
Veröffentlicht: (2025)
von: Bian, Yutong, et al.
Veröffentlicht: (2025)
o1-Coder: an o1 Replication for Coding
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
MIMIC-Py: An Extensible Tool for Personality-Driven Automated Game Testing with Large Language Models
von: Chen, Yifei, et al.
Veröffentlicht: (2026)
von: Chen, Yifei, et al.
Veröffentlicht: (2026)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
von: Hong, Sirui, et al.
Veröffentlicht: (2026)
von: Hong, Sirui, et al.
Veröffentlicht: (2026)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
GUITester: Enabling GUI Agents for Exploratory Defect Discovery
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
Fragmented Layer Grouping in GUI Designs Through Graph Learning Based on Multimodal Information
von: Chen, Yunnong, et al.
Veröffentlicht: (2024)
von: Chen, Yunnong, et al.
Veröffentlicht: (2024)
Demystifying Issues, Causes and Solutions in LLM Open-Source Projects
von: Cai, Yangxiao, et al.
Veröffentlicht: (2024)
von: Cai, Yangxiao, et al.
Veröffentlicht: (2024)
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
von: Arrieta, Aitor, et al.
Veröffentlicht: (2025)
von: Arrieta, Aitor, et al.
Veröffentlicht: (2025)
Qualitative Evaluation of LLM-Designed GUI
von: Sawicki, Bartosz, et al.
Veröffentlicht: (2026)
von: Sawicki, Bartosz, et al.
Veröffentlicht: (2026)
LLMParser: An Exploratory Study on Using Large Language Models for Log Parsing
von: Ma, Zeyang, et al.
Veröffentlicht: (2024)
von: Ma, Zeyang, et al.
Veröffentlicht: (2024)
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
von: Han, Qijun, et al.
Veröffentlicht: (2026)
von: Han, Qijun, et al.
Veröffentlicht: (2026)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
von: Cui, Yi
Veröffentlicht: (2025)
von: Cui, Yi
Veröffentlicht: (2025)
Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
Just-in-Time Catching Test Generation at Meta
von: Becker, Matthew, et al.
Veröffentlicht: (2026)
von: Becker, Matthew, et al.
Veröffentlicht: (2026)
SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration
von: Wang, Qingni, et al.
Veröffentlicht: (2026)
von: Wang, Qingni, et al.
Veröffentlicht: (2026)
AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
von: Leung, Ho Fai, et al.
Veröffentlicht: (2025)
von: Leung, Ho Fai, et al.
Veröffentlicht: (2025)
Towards a Framework for Openness in Foundation Models: Proceedings from the Columbia Convening on Openness in Artificial Intelligence
von: Basdevant, Adrien, et al.
Veröffentlicht: (2024)
von: Basdevant, Adrien, et al.
Veröffentlicht: (2024)
Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
von: Huang, Yuheng, et al.
Veröffentlicht: (2023)
von: Huang, Yuheng, et al.
Veröffentlicht: (2023)
An Empirical Study of OpenAI API Discussions on Stack Overflow
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
Clarifying Semantics of In-Context Examples for Unit Test Generation
von: Yang, Chen, et al.
Veröffentlicht: (2025)
von: Yang, Chen, et al.
Veröffentlicht: (2025)
LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming
von: Sung, Sicheol, et al.
Veröffentlicht: (2025)
von: Sung, Sicheol, et al.
Veröffentlicht: (2025)
Evaluating Human Trajectory Prediction with Metamorphic Testing
von: Spieker, Helge, et al.
Veröffentlicht: (2024)
von: Spieker, Helge, et al.
Veröffentlicht: (2024)
Towards Better Correctness and Efficiency in Code Generation
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
LLM-Based Test Case Generation in DBMS through Monte Carlo Tree Search
von: Chen, Yujia, et al.
Veröffentlicht: (2026)
von: Chen, Yujia, et al.
Veröffentlicht: (2026)
ReCatcher: Towards LLMs Regression Testing for Code Generation
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2025)
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2025)
InfraMind: A Novel Exploration-based GUI Agentic Framework for Mission-critical Industrial Management
von: Lin, Liangtao, et al.
Veröffentlicht: (2025)
von: Lin, Liangtao, et al.
Veröffentlicht: (2025)
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
von: Xu, Jingxuan, et al.
Veröffentlicht: (2025)
von: Xu, Jingxuan, et al.
Veröffentlicht: (2025)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
Towards Reliable LLM-Driven Fuzz Testing: Vision and Road Ahead
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges
von: Harman, Mark, et al.
Veröffentlicht: (2025)
von: Harman, Mark, et al.
Veröffentlicht: (2025)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
Evaluating LLM-Based Test Generation Under Software Evolution
von: Haroon, Sabaat, et al.
Veröffentlicht: (2026)
von: Haroon, Sabaat, et al.
Veröffentlicht: (2026)
Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?
von: Qian, Yi, et al.
Veröffentlicht: (2026)
von: Qian, Yi, et al.
Veröffentlicht: (2026)
Towards Secure Program Partitioning for Smart Contracts with LLM's In-Context Learning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
SELF-REDRAFT: Eliciting Intrinsic Exploration-Exploitation Balance in Test-Time Scaling for Code Generation
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
von: Xu, Xinbo, et al.
Veröffentlicht: (2026)
von: Xu, Xinbo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
von: Bian, Yutong, et al.
Veröffentlicht: (2025) -
o1-Coder: an o1 Replication for Coding
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024) -
MIMIC-Py: An Extensible Tool for Personality-Driven Automated Game Testing with Large Language Models
von: Chen, Yifei, et al.
Veröffentlicht: (2026) -
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
von: Hong, Sirui, et al.
Veröffentlicht: (2026) -
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)