SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Syed Yusuf, Feng, Shiwei, Bae, Chanwoo, Zhang, Calix Barrus Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
by: Wu, Jiaqing, et al.
Published: (2026)
by: Wu, Jiaqing, et al.
Published: (2026)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
by: Bhardwaj, Varun Pratap
Published: (2026)
by: Bhardwaj, Varun Pratap
Published: (2026)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024)
by: Li, Meiziniu, et al.
Published: (2024)
RepoLaunch: Automating Build&Test Pipeline of Code Repositories on ANY Language and ANY Platform
by: Li, Kenan, et al.
Published: (2026)
by: Li, Kenan, et al.
Published: (2026)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022)
by: Li, Meiziniu, et al.
Published: (2022)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
by: Liu, Zhuoyao, et al.
Published: (2026)
by: Liu, Zhuoyao, et al.
Published: (2026)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
by: Bradbury, Jeremy S., et al.
Published: (2024)
by: Bradbury, Jeremy S., et al.
Published: (2024)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
by: Xu, Qiyuan, et al.
Published: (2026)
by: Xu, Qiyuan, et al.
Published: (2026)
AutoBridge: Automating Smart Device Integration with Centralized Platform
by: Liu, Siyuan, et al.
Published: (2025)
by: Liu, Siyuan, et al.
Published: (2025)
CodeEvolve: LLM-Driven Evolutionary Optimization with Runtime-Enriched Target Selection for Multi-Language Code Enhancement
by: Borra, Ajay Krishna, et al.
Published: (2026)
by: Borra, Ajay Krishna, et al.
Published: (2026)
Review Beats Planning: Dual-Model Interaction Patterns for Code Synthesis
by: Miller, Jan
Published: (2026)
by: Miller, Jan
Published: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
by: Li, Kenan, et al.
Published: (2026)
by: Li, Kenan, et al.
Published: (2026)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
by: Zietsman, Christo
Published: (2026)
by: Zietsman, Christo
Published: (2026)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
by: Terragni, Valerio
Published: (2026)
by: Terragni, Valerio
Published: (2026)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
by: Ramírez, Aurora, et al.
Published: (2024)
by: Ramírez, Aurora, et al.
Published: (2024)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
by: Li, Meiziniu, et al.
Published: (2026)
by: Li, Meiziniu, et al.
Published: (2026)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026)
by: Ravi, Ravin, et al.
Published: (2026)
Automated structural testing of LLM-based agents: methods, framework, and case studies
by: Kohl, Jens, et al.
Published: (2026)
by: Kohl, Jens, et al.
Published: (2026)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
LLMORPH: Automated Metamorphic Testing of Large Language Models
by: Cho, Steven, et al.
Published: (2026)
by: Cho, Steven, et al.
Published: (2026)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
by: Chang, Hung-Fu, et al.
Published: (2025)
by: Chang, Hung-Fu, et al.
Published: (2025)
TDAD: Test-Driven Agentic Development - Reducing Code Regressions in AI Coding Agents via Graph-Based Impact Analysis
by: Alonso, Pepe, et al.
Published: (2026)
by: Alonso, Pepe, et al.
Published: (2026)
MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
by: Sartaj, Hassan, et al.
Published: (2024)
by: Sartaj, Hassan, et al.
Published: (2024)
Understanding and Detecting Flaky Builds in GitHub Actions
by: Ge, Wenhao, et al.
Published: (2026)
by: Ge, Wenhao, et al.
Published: (2026)
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
by: Meng, Haoming
Published: (2026)
by: Meng, Haoming
Published: (2026)
Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches
by: Wienczkowski, Michael
Published: (2026)
by: Wienczkowski, Michael
Published: (2026)
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)
by: Gunnu, Sravani, et al.
Published: (2025)
Experience with GitHub Copilot for Developer Productivity at Zoominfo
by: Bakal, Gal, et al.
Published: (2025)
by: Bakal, Gal, et al.
Published: (2025)
Multi-Agent Code Verification via Information Theory
by: Rajan, Shreshth
Published: (2025)
by: Rajan, Shreshth
Published: (2025)
An LSTM-based Test Selection Method for Self-Driving Cars
by: Güllü, Ali, et al.
Published: (2025)
by: Güllü, Ali, et al.
Published: (2025)
Automated Code Fix Suggestions for Accessibility Issues in Mobile Apps
by: Mehralian, Forough, et al.
Published: (2024)
by: Mehralian, Forough, et al.
Published: (2024)
RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing
by: Guo, Jinyao, et al.
Published: (2025)
by: Guo, Jinyao, et al.
Published: (2025)
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
by: Cambronero, José, et al.
Published: (2025)
by: Cambronero, José, et al.
Published: (2025)
Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation
by: Downing, Mara, et al.
Published: (2025)
by: Downing, Mara, et al.
Published: (2025)
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
Similar Items
-
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
by: Wu, Jiaqing, et al.
Published: (2026) -
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
by: Bhardwaj, Varun Pratap
Published: (2026) -
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026) -
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025) -
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)