"Are We Done Yet?": A Vision-Based Judge for Autonomous Task Completion of Computer Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Sumyk, Marta, Kosovan, Oleksandr |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents
by: Sumyk, Marta, et al.
Published: (2026)
by: Sumyk, Marta, et al.
Published: (2026)
Toward Agentic RAG for Ukrainian
by: Sumyk, Marta, et al.
Published: (2026)
by: Sumyk, Marta, et al.
Published: (2026)
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation
by: Muryn, Viktor, et al.
Published: (2025)
by: Muryn, Viktor, et al.
Published: (2025)
ParaView-MCP: An Autonomous Visualization Agent with Direct Tool Use
by: Liu, Shusen, et al.
Published: (2025)
by: Liu, Shusen, et al.
Published: (2025)
What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
by: Ma, Qianou, et al.
Published: (2024)
by: Ma, Qianou, et al.
Published: (2024)
Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
AgentLens: Visual Analysis for Agent Behaviors in LLM-based Autonomous Systems
by: Lu, Jiaying, et al.
Published: (2024)
by: Lu, Jiaying, et al.
Published: (2024)
Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices
by: Yan, Xiao, et al.
Published: (2025)
by: Yan, Xiao, et al.
Published: (2025)
LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
IntentCUA: Learning Intent-level Representations for Skill Abstraction and Multi-Agent Planning in Computer-Use Agents
by: Lee, Seoyoung, et al.
Published: (2026)
by: Lee, Seoyoung, et al.
Published: (2026)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)
by: Natalie, Rosiana, et al.
Published: (2025)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
Learning to Assign Prediction Tasks to Agents with Capacity Constraints
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
Scenarios in Computing Research: A Systematic Review of the Use of Scenario Methods for Exploring the Future of Computing Technologies in Society
by: Barnett, Julia, et al.
Published: (2025)
by: Barnett, Julia, et al.
Published: (2025)
A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions
by: Sager, Pascal J., et al.
Published: (2025)
by: Sager, Pascal J., et al.
Published: (2025)
Explainable AI Enhances Glaucoma Referrals, Yet the Human-AI Team Still Falls Short of the AI Alone
by: Gomez, Catalina, et al.
Published: (2024)
by: Gomez, Catalina, et al.
Published: (2024)
MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
by: Kong, Yi, et al.
Published: (2025)
by: Kong, Yi, et al.
Published: (2025)
See or Recall: A Sanity Check for the Role of Vision in Solving Visualization Question Answer Tasks with Multimodal LLMs
by: Li, Zhimin, et al.
Published: (2025)
by: Li, Zhimin, et al.
Published: (2025)
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024)
by: Niu, Runliang, et al.
Published: (2024)
HTN-Based Tutors: A New Intelligent Tutoring Framework Based on Hierarchical Task Networks
by: Siddiqui, Momin N., et al.
Published: (2024)
by: Siddiqui, Momin N., et al.
Published: (2024)
CAAP: Context-Aware Action Planning Prompting to Solve Computer Tasks with Front-End UI Only
by: Cho, Junhee, et al.
Published: (2024)
by: Cho, Junhee, et al.
Published: (2024)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
by: Huq, Faria, et al.
Published: (2025)
by: Huq, Faria, et al.
Published: (2025)
Anticipating User Needs: Insights from Design Fiction on Conversational Agents for Computational Thinking
by: Penney, Jacob, et al.
Published: (2023)
by: Penney, Jacob, et al.
Published: (2023)
Building Persona-Based Agents On Demand: Tailoring Multi-Agent Workflows to User Needs
by: Arbore, Giuseppe, et al.
Published: (2026)
by: Arbore, Giuseppe, et al.
Published: (2026)
ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
Cases of EFL Secondary Students' Prompt Engineering Pathways to Complete a Writing Task with ChatGPT
by: Woo, David James, et al.
Published: (2023)
by: Woo, David James, et al.
Published: (2023)
Creating General User Models from Computer Use
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
AgentEconomist: An End-to-end Agentic System Translating Economic Intuitions into Executable Computational Experiments
by: Chen, Jiaju, et al.
Published: (2026)
by: Chen, Jiaju, et al.
Published: (2026)
MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
by: Luera, Reuben A., et al.
Published: (2025)
by: Luera, Reuben A., et al.
Published: (2025)
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
The HCI GenAI CO2ST Calculator: A Tool for Calculating the Carbon Footprint of Generative AI Use in Human-Computer Interaction Research
by: Inie, Nanna, et al.
Published: (2025)
by: Inie, Nanna, et al.
Published: (2025)
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
by: Rahman, Hasibur, et al.
Published: (2025)
by: Rahman, Hasibur, et al.
Published: (2025)
Longitudinal Study on Social and Emotional Use of AI Conversational Agent
by: Chandra, Mohit, et al.
Published: (2025)
by: Chandra, Mohit, et al.
Published: (2025)
A Scoping Review of the Ethical Perspectives on Anthropomorphising Large Language Model-Based Conversational Agents
by: Ferrario, Andrea, et al.
Published: (2026)
by: Ferrario, Andrea, et al.
Published: (2026)
Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code Completions
by: Vasconcelos, Helena, et al.
Published: (2023)
by: Vasconcelos, Helena, et al.
Published: (2023)
TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks
by: Yin, Yiwen, et al.
Published: (2025)
by: Yin, Yiwen, et al.
Published: (2025)
Effective Yet Ephemeral Propaganda Defense: There Needs to Be More than One-Shot Inoculation to Enhance Critical Thinking
by: Hoferer, Nicolas, et al.
Published: (2025)
by: Hoferer, Nicolas, et al.
Published: (2025)
Analysis of AI Effectiveness in Reducing Human Errors in Processing Transportation Requests
by: Korostin, Oleksandr
Published: (2025)
by: Korostin, Oleksandr
Published: (2025)
Reactive Writers: How Co-Writing with AI Changes How We Engage with Ideas
by: Bhat, Advait, et al.
Published: (2026)
by: Bhat, Advait, et al.
Published: (2026)
Similar Items
-
CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents
by: Sumyk, Marta, et al.
Published: (2026) -
Toward Agentic RAG for Ukrainian
by: Sumyk, Marta, et al.
Published: (2026) -
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation
by: Muryn, Viktor, et al.
Published: (2025) -
ParaView-MCP: An Autonomous Visualization Agent with Direct Tool Use
by: Liu, Shusen, et al.
Published: (2025) -
What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
by: Ma, Qianou, et al.
Published: (2024)