OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Jinbiao, Ma, Qianran, Zhao, Yilun, Zhou, Xiao, Ni, Kangqi, Gan, Guo, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Step-level Optimization for Efficient Computer-use Agents
by: Wei, Jinbiao, et al.
Published: (2026)
by: Wei, Jinbiao, et al.
Published: (2026)
ANCHOR: Branch-Point Data Generation for GUI Agents
by: Wei, Jinbiao, et al.
Published: (2026)
by: Wei, Jinbiao, et al.
Published: (2026)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
Computer Use at the Edge of the Statistical Precipice
by: D'Oro, Pierluca, et al.
Published: (2026)
by: D'Oro, Pierluca, et al.
Published: (2026)
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering
by: Guo, Chuanzhe, et al.
Published: (2026)
by: Guo, Chuanzhe, et al.
Published: (2026)
REVERE: Reflective Evolving Research Engineer for Scientific Workflows
by: Gangireddi, Balaji Dinesh, et al.
Published: (2026)
by: Gangireddi, Balaji Dinesh, et al.
Published: (2026)
Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute
by: Ma, Yingwei, et al.
Published: (2025)
by: Ma, Yingwei, et al.
Published: (2025)
Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
by: Wang, Xingyao, et al.
Published: (2025)
by: Wang, Xingyao, et al.
Published: (2025)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
by: Chen, Mouxiang, et al.
Published: (2026)
by: Chen, Mouxiang, et al.
Published: (2026)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
by: Riddell, Martin, et al.
Published: (2024)
by: Riddell, Martin, et al.
Published: (2024)
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software
by: Meng, Lingkai, et al.
Published: (2025)
by: Meng, Lingkai, et al.
Published: (2025)
LocAgent: Graph-Guided LLM Agents for Code Localization
by: Chen, Zhaoling, et al.
Published: (2025)
by: Chen, Zhaoling, et al.
Published: (2025)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
by: Zhang, Zehua, et al.
Published: (2025)
by: Zhang, Zehua, et al.
Published: (2025)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
by: Han, Tingxu, et al.
Published: (2026)
by: Han, Tingxu, et al.
Published: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
by: Wang, Xingyao, et al.
Published: (2024)
by: Wang, Xingyao, et al.
Published: (2024)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
by: Trae Research Team, et al.
Published: (2025)
by: Trae Research Team, et al.
Published: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Verifying LLM-Generated Code in the Context of Software Verification with Ada/SPARK
by: Cramer, Marcos, et al.
Published: (2025)
by: Cramer, Marcos, et al.
Published: (2025)
Leveraging Large Language Models for Code Translation and Software Development in Scientific Computing
by: Dhruv, Akash, et al.
Published: (2024)
by: Dhruv, Akash, et al.
Published: (2024)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
by: Xiao, Yijia, et al.
Published: (2025)
by: Xiao, Yijia, et al.
Published: (2025)
AgentHub: A Registry for Discoverable, Verifiable, and Reproducible AI Agents
by: Pautsch, Erik, et al.
Published: (2025)
by: Pautsch, Erik, et al.
Published: (2025)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
ArchAgent: Scalable Legacy Software Architecture Recovery with LLMs
by: Pan, Rusheng, et al.
Published: (2026)
by: Pan, Rusheng, et al.
Published: (2026)
Same Signal, Different Semantics: A Cross-Framework Behavioral Analysis of Software Engineering Agents
by: Ma, Wei, et al.
Published: (2026)
by: Ma, Wei, et al.
Published: (2026)
Open-Source AI-based SE Tools: Opportunities and Challenges of Collaborative Software Learning
by: Lin, Zhihao, et al.
Published: (2024)
by: Lin, Zhihao, et al.
Published: (2024)
Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents
by: Zhao, Chenyu, et al.
Published: (2026)
by: Zhao, Chenyu, et al.
Published: (2026)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
by: Yu, Zhongming, et al.
Published: (2025)
by: Yu, Zhongming, et al.
Published: (2025)
Unified Software Engineering Agent as AI Software Engineer
by: Applis, Leonhard, et al.
Published: (2025)
by: Applis, Leonhard, et al.
Published: (2025)
Self-Evolving Software Agents
by: Robol, Marco, et al.
Published: (2026)
by: Robol, Marco, et al.
Published: (2026)
TOM-SWE: User Mental Modeling For Software Engineering Agents
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
AgentStepper: Interactive Debugging of Software Development Agents
by: Hutter, Robert, et al.
Published: (2026)
by: Hutter, Robert, et al.
Published: (2026)
Toward a Science of Intent: Closure Gaps and Delegation Envelopes for Open-World AI Agents
by: Armesto, Maximiliano, et al.
Published: (2026)
by: Armesto, Maximiliano, et al.
Published: (2026)
Towards Verifiably Safe Tool Use for LLM Agents
by: Doshi, Aarya, et al.
Published: (2026)
by: Doshi, Aarya, et al.
Published: (2026)
Towards Adaptive Software Agents for Debugging
by: Majdoub, Yacine, et al.
Published: (2025)
by: Majdoub, Yacine, et al.
Published: (2025)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
by: Pandey, Ruchika, et al.
Published: (2024)
by: Pandey, Ruchika, et al.
Published: (2024)
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
by: Yu, Guangba, et al.
Published: (2026)
by: Yu, Guangba, et al.
Published: (2026)
From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents
by: Ma, Murong, et al.
Published: (2026)
by: Ma, Murong, et al.
Published: (2026)
Similar Items
-
Step-level Optimization for Efficient Computer-use Agents
by: Wei, Jinbiao, et al.
Published: (2026) -
ANCHOR: Branch-Point Data Generation for GUI Agents
by: Wei, Jinbiao, et al.
Published: (2026) -
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024) -
Computer Use at the Edge of the Statistical Precipice
by: D'Oro, Pierluca, et al.
Published: (2026) -
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering
by: Guo, Chuanzhe, et al.
Published: (2026)