UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Nayak, Shravan, Jian, Xiangru, Lin, Kevin Qinghong, Rodriguez, Juan A., Kalsi, Montek, Awal, Rabiul, Chapados, Nicolas, Özsu, M. Tamer, Agrawal, Aishwarya, Vazquez, David, Pal, Christopher, Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025)
by: Rodriguez, Juan A., et al.
Published: (2025)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023)
by: Zhang, Le, et al.
Published: (2023)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
by: Wang, Xiaoqiang, et al.
Published: (2026)
by: Wang, Xiaoqiang, et al.
Published: (2026)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
Benchmarking Vision Language Models for Cultural Understanding
by: Nayak, Shravan, et al.
Published: (2024)
by: Nayak, Shravan, et al.
Published: (2024)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
InteracSPARQL: An Interactive System for SPARQL Query Refinement Using Natural Language Explanations
by: Jian, Xiangru, et al.
Published: (2025)
by: Jian, Xiangru, et al.
Published: (2025)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
The Promise of RL for Autoregressive Image Editing
by: Ahmadi, Saba, et al.
Published: (2025)
by: Ahmadi, Saba, et al.
Published: (2025)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Foundations and Scoping of Data Science
by: Özsu, M. Tamer
Published: (2023)
by: Özsu, M. Tamer
Published: (2023)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
Reinforcement Learning from Delayed Observations via World Models
by: Karamzade, Armin, et al.
Published: (2024)
by: Karamzade, Armin, et al.
Published: (2024)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
by: Rodriguez, Juan, et al.
Published: (2024)
by: Rodriguez, Juan, et al.
Published: (2024)
Principles of distributed database systems / M. Tamer Özsu, Patrick Valduriez
by: aÖzsu, M. Tamer
Published: (1999)
by: aÖzsu, M. Tamer
Published: (1999)
Principles of distributed database systems / M. Tamer Ôzsu, Patrick Valduriez
by: aÔzsu, M. Tamer
Published: (2011)
by: aÔzsu, M. Tamer
Published: (2011)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025)
by: Hu, Siyuan, et al.
Published: (2025)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025)
by: Maheshwary, Rishabh, et al.
Published: (2025)
LazyVLM: Neuro-Symbolic Approach to Video Analytics
by: Jian, Xiangru, et al.
Published: (2025)
by: Jian, Xiangru, et al.
Published: (2025)
A Demonstration of SQLyzr: A Platform for Fine-Grained Text-to-SQL Evaluation and Analysis
by: Abedini, Sepideh, et al.
Published: (2026)
by: Abedini, Sepideh, et al.
Published: (2026)
Transient Concepts in Streaming Graphs
by: Sheshbolouki, Aida, et al.
Published: (2025)
by: Sheshbolouki, Aida, et al.
Published: (2025)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
by: Xu, Wenpeng
Published: (2026)
by: Xu, Wenpeng
Published: (2026)
Discovering Failure Modes in Vision-Language Models using RL
by: Jain, Kanishk, et al.
Published: (2026)
by: Jain, Kanishk, et al.
Published: (2026)
Low-Latency Sliding Window Connectivity
by: Zhang, Chao, et al.
Published: (2024)
by: Zhang, Chao, et al.
Published: (2024)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
Similar Items
-
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025) -
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025) -
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025) -
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026) -
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)