Toward Autonomous UI Exploration: The UIExplorer Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Nica, Andrei Cristian, Shanbhogue, Akshaya Vishnu Kudlu, Shah, Harshil, Cambray, Aleix, Berariu, Tudor, Maystre, Lucas, Barber, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025)
by: Dolga, Rares, et al.
Published: (2025)
ComfySearch: Autonomous Exploration and Reasoning for ComfyUI Workflows
by: Su, Jinwei, et al.
Published: (2026)
by: Su, Jinwei, et al.
Published: (2026)
Optimizing Audio Recommendations for the Long-Term: A Reinforcement Learning Perspective
by: Maystre, Lucas, et al.
Published: (2023)
by: Maystre, Lucas, et al.
Published: (2023)
NAAMSE: Framework for Evolutionary Security Evaluation of Agents
by: Pai, Kunal, et al.
Published: (2026)
by: Pai, Kunal, et al.
Published: (2026)
Teaching by Failure: Counter-Example-Driven Curricula for Transformer Self-Improvement
by: Vejendla, Harshil
Published: (2025)
by: Vejendla, Harshil
Published: (2025)
RewriteNets: End-to-End Trainable String-Rewriting for Generative Sequence Modeling
by: Vejendla, Harshil
Published: (2026)
by: Vejendla, Harshil
Published: (2026)
Generative UI: LLMs are Effective UI Generators
by: Leviathan, Yaniv, et al.
Published: (2026)
by: Leviathan, Yaniv, et al.
Published: (2026)
SliceMoE: Routing Embedding Slices Instead of Tokens for Fine-Grained and Balanced Transformer Scaling
by: Vejendla, Harshil
Published: (2025)
by: Vejendla, Harshil
Published: (2025)
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
by: Chai, Yuxiang, et al.
Published: (2026)
by: Chai, Yuxiang, et al.
Published: (2026)
When Embedding Models Meet: Procrustes Bounds and Applications
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
by: Garkot, Sofiya, et al.
Published: (2025)
by: Garkot, Sofiya, et al.
Published: (2025)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
by: Xue, Xiangyuan, et al.
Published: (2024)
by: Xue, Xiangyuan, et al.
Published: (2024)
Impatient Bandits: Optimizing for the Long-Term Without Delay
by: Zhang, Kelly W., et al.
Published: (2025)
by: Zhang, Kelly W., et al.
Published: (2025)
Nomad: Autonomous Exploration and Discovery
by: Jia, Bokang, et al.
Published: (2026)
by: Jia, Bokang, et al.
Published: (2026)
Magentic-UI: Towards Human-in-the-loop Agentic Systems
by: Mozannar, Hussein, et al.
Published: (2025)
by: Mozannar, Hussein, et al.
Published: (2025)
ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark
by: Ciobanu, Ioan-Paul, et al.
Published: (2025)
by: Ciobanu, Ioan-Paul, et al.
Published: (2025)
Proximal Policy Optimization with Adaptive Exploration
by: Lixandru, Andrei
Published: (2024)
by: Lixandru, Andrei
Published: (2024)
Benchmarking LLMs for Predictive Applications in the Intensive Care Units
by: Malhotra, Chehak, et al.
Published: (2025)
by: Malhotra, Chehak, et al.
Published: (2025)
Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
GhostUI: Unveiling Hidden Interactions in Mobile UI
by: Kweon, Minkyu, et al.
Published: (2026)
by: Kweon, Minkyu, et al.
Published: (2026)
Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
by: Polonuer, Joaquín, et al.
Published: (2026)
by: Polonuer, Joaquín, et al.
Published: (2026)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
MAIC-UI: Making Interactive Courseware with Generative UI
by: Tu, Shangqing, et al.
Published: (2026)
by: Tu, Shangqing, et al.
Published: (2026)
MANA: Towards Efficient Mobile Ad Detection via Multimodal Agentic UI Navigation
by: Zhao, Yizhe, et al.
Published: (2026)
by: Zhao, Yizhe, et al.
Published: (2026)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
by: Jarca, Andrei, et al.
Published: (2025)
by: Jarca, Andrei, et al.
Published: (2025)
Towards Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-critical Scenarios
by: Li, Jingzheng, et al.
Published: (2025)
by: Li, Jingzheng, et al.
Published: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
by: Xu, Qiao, et al.
Published: (2026)
by: Xu, Qiao, et al.
Published: (2026)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
by: Luera, Reuben A., et al.
Published: (2025)
by: Luera, Reuben A., et al.
Published: (2025)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
AutoGameUI: Constructing High-Fidelity GameUI via Multimodal Correspondence Matching
by: Tang, Zhongliang, et al.
Published: (2024)
by: Tang, Zhongliang, et al.
Published: (2024)
Rip Current Segmentation: A Novel Benchmark and YOLOv8 Baseline Results
by: Dumitriu, Andrei, et al.
Published: (2025)
by: Dumitriu, Andrei, et al.
Published: (2025)
Autonomous Implicit Indoor Scene Reconstruction with Frontier Exploration
by: Zeng, Jing, et al.
Published: (2024)
by: Zeng, Jing, et al.
Published: (2024)
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
SE(3)-Stochastic Flow Matching for Protein Backbone Generation
by: Bose, Avishek Joey, et al.
Published: (2023)
by: Bose, Avishek Joey, et al.
Published: (2023)
A Rosetta Stone for AI Benchmarks
by: Ho, Anson, et al.
Published: (2025)
by: Ho, Anson, et al.
Published: (2025)
Similar Items
-
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025) -
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025) -
ComfySearch: Autonomous Exploration and Reasoning for ComfyUI Workflows
by: Su, Jinwei, et al.
Published: (2026) -
Optimizing Audio Recommendations for the Long-Term: A Reinforcement Learning Perspective
by: Maystre, Lucas, et al.
Published: (2023) -
NAAMSE: Framework for Evolutionary Security Evaluation of Agents
by: Pai, Kunal, et al.
Published: (2026)