GPA: Learning GUI Process Automation from Demonstrations
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zirui, Liew, Jun Hao, Yang, Yan, Yang, Wenzhuo, Luo, Ziyang, Sahoo, Doyen, Savarese, Silvio, Li, Junnan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Reinforcement Learning for Automated Web GUI Testing
by: Gu, Zhiyu, et al.
Published: (2025)
by: Gu, Zhiyu, et al.
Published: (2025)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
GUing: A Mobile GUI Search Engine using a Vision-Language Model
by: Wei, Jialiang, et al.
Published: (2024)
by: Wei, Jialiang, et al.
Published: (2024)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
Natural Adversaries: Fuzzing Autonomous Vehicles with Realistic Roadside Object Placements
by: Sun, Yang, et al.
Published: (2024)
by: Sun, Yang, et al.
Published: (2024)
GUI Agents for Continual Game Generation
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
Benchmarking Image Perturbations for Testing Automated Driving Assistance Systems
by: Lambertenghi, Stefano Carlo, et al.
Published: (2025)
by: Lambertenghi, Stefano Carlo, et al.
Published: (2025)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
by: Li, Kaixin, et al.
Published: (2024)
by: Li, Kaixin, et al.
Published: (2024)
SVRepair: Structured Visual Reasoning for Automated Program Repair
by: Tang, Xiaoxuan, et al.
Published: (2026)
by: Tang, Xiaoxuan, et al.
Published: (2026)
Temac: Multi-Agent Collaboration for Automated Web GUI Testing
by: Liu, Chenxu, et al.
Published: (2025)
by: Liu, Chenxu, et al.
Published: (2025)
Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts
by: Lu, Xiao, et al.
Published: (2026)
by: Lu, Xiao, et al.
Published: (2026)
Self-Elicitation of Requirements with Automated GUI Prototyping
by: Kolthoff, Kristian, et al.
Published: (2024)
by: Kolthoff, Kristian, et al.
Published: (2024)
Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?
by: Lyu, Yunbo, et al.
Published: (2025)
by: Lyu, Yunbo, et al.
Published: (2025)
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
by: Meng, Fanqing, et al.
Published: (2026)
by: Meng, Fanqing, et al.
Published: (2026)
How Smart Is Your GUI Agent? A Framework for the Future of Software Interaction
by: Feng, Sidong, et al.
Published: (2026)
by: Feng, Sidong, et al.
Published: (2026)
Ear-Keeper: A Cross-Platform AI System for Rapid and Accurate Ear Disease Diagnosis
by: Lu, Feiyan, et al.
Published: (2023)
by: Lu, Feiyan, et al.
Published: (2023)
AwesomeMeta+: A Mixed-Prototyping Meta-Learning System Supporting AI Application Design Anywhere
by: Wang, Jingyao, et al.
Published: (2023)
by: Wang, Jingyao, et al.
Published: (2023)
GUIDE: LLM-Driven GUI Generation Decomposition for Automated Prototyping
by: Kolthoff, Kristian, et al.
Published: (2025)
by: Kolthoff, Kristian, et al.
Published: (2025)
On-Demand Scenario Generation for Testing Automated Driving Systems
by: Yan, Songyang, et al.
Published: (2025)
by: Yan, Songyang, et al.
Published: (2025)
ROMAN: Reward-Orchestrated Multi-Head Attention Network for Autonomous Driving System Testing
by: Chi, Jianlei, et al.
Published: (2026)
by: Chi, Jianlei, et al.
Published: (2026)
WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
by: Liang, Shanchao, et al.
Published: (2024)
by: Liang, Shanchao, et al.
Published: (2024)
PopSweeper: Automatically Detecting and Resolving App-Blocking Pop-Ups to Assist Automated Mobile GUI Testing
by: Guo, Linqiang, et al.
Published: (2024)
by: Guo, Linqiang, et al.
Published: (2024)
Cross-Breed Pig Identification Using Auricular Vein Pattern Recognition: A Machine Learning Approach for Small-Scale Farming Applications
by: Nsengiyumvaa, Emmanuel, et al.
Published: (2025)
by: Nsengiyumvaa, Emmanuel, et al.
Published: (2025)
Quality at the Tail of Machine Learning Inference
by: Yang, Zhengxin, et al.
Published: (2022)
by: Yang, Zhengxin, et al.
Published: (2022)
Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression
by: Zhong, Jianping, et al.
Published: (2026)
by: Zhong, Jianping, et al.
Published: (2026)
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
by: Xu, Zhenran, et al.
Published: (2025)
by: Xu, Zhenran, et al.
Published: (2025)
FuncDroid: Towards Inter-Functional Flows for Comprehensive Mobile App GUI Testing
by: He, Jinlong, et al.
Published: (2026)
by: He, Jinlong, et al.
Published: (2026)
EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering
by: Dou, Zhifei, et al.
Published: (2026)
by: Dou, Zhifei, et al.
Published: (2026)
MaCTG: Multi-Agent Collaborative Thought Graph for Automatic Programming
by: Zhao, Zixiao, et al.
Published: (2024)
by: Zhao, Zixiao, et al.
Published: (2024)
GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes
by: Kolthoff, Kristian, et al.
Published: (2025)
by: Kolthoff, Kristian, et al.
Published: (2025)
From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation
by: Xie, Yuhang, et al.
Published: (2025)
by: Xie, Yuhang, et al.
Published: (2025)
GUI Element Detection Using SOTA YOLO Deep Learning Models
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
SWAN -- Enabling Fast and Mobile Histopathology Image Annotation through Swipeable Interfaces
by: Banerjee, Sweta, et al.
Published: (2025)
by: Banerjee, Sweta, et al.
Published: (2025)
DD-CAM: Minimal Sufficient Explanations for Vision Models Using Delta Debugging
by: Khadka, Krishna, et al.
Published: (2026)
by: Khadka, Krishna, et al.
Published: (2026)
Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality
by: Hong, Danfeng, et al.
Published: (2026)
by: Hong, Danfeng, et al.
Published: (2026)
Technical Report for Argoverse2 Scenario Mining Challenges on Iterative Error Correction and Spatially-Aware Prompting
by: Chen, Yifei, et al.
Published: (2025)
by: Chen, Yifei, et al.
Published: (2025)
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
by: Lin, Haojia, et al.
Published: (2025)
by: Lin, Haojia, et al.
Published: (2025)
How Far Can VLMs Go for Visual Bug Detection? Studying 19,738 Keyframes from 41 Hours of Gameplay Videos
by: Lu, Wentao, et al.
Published: (2026)
by: Lu, Wentao, et al.
Published: (2026)
Earth Embeddings as Products: Taxonomy, Ecosystem, and Standardized Access
by: Fang, Heng, et al.
Published: (2026)
by: Fang, Heng, et al.
Published: (2026)
Interpretable Gallbladder Ultrasound Diagnosis: A Lightweight Web-Mobile Software Platform with Real-Time XAI
by: Bhoyan, Fuyad Hasan, et al.
Published: (2025)
by: Bhoyan, Fuyad Hasan, et al.
Published: (2025)
Similar Items
-
Deep Reinforcement Learning for Automated Web GUI Testing
by: Gu, Zhiyu, et al.
Published: (2025) -
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024) -
GUing: A Mobile GUI Search Engine using a Vision-Language Model
by: Wei, Jialiang, et al.
Published: (2024) -
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024) -
Natural Adversaries: Fuzzing Autonomous Vehicles with Realistic Roadside Object Placements
by: Sun, Yang, et al.
Published: (2024)