Programming with Pixels: Can Computer-Use Agents do Software Engineering?
Fuente:
arXiv
Saved in:
| Main Authors: | Aggarwal, Pranjal, Welleck, Sean |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
miniCodeProps: a Minimal Benchmark for Proving Code Properties
by: Lohn, Evan, et al.
Published: (2024)
by: Lohn, Evan, et al.
Published: (2024)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
Assessing the Use of AutoML for Data-Driven Software Engineering
by: Calefato, Fabio, et al.
Published: (2023)
by: Calefato, Fabio, et al.
Published: (2023)
Agint: Agentic Graph Compilation for Software Engineering Agents
by: Chivukula, Abhi, et al.
Published: (2025)
by: Chivukula, Abhi, et al.
Published: (2025)
A Systematic Literature Review on the Use of Machine Learning in Software Engineering
by: Fred, Nyaga, et al.
Published: (2024)
by: Fred, Nyaga, et al.
Published: (2024)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
by: Bula, Timothy, et al.
Published: (2025)
by: Bula, Timothy, et al.
Published: (2025)
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
by: Miserendino, Samuel, et al.
Published: (2025)
by: Miserendino, Samuel, et al.
Published: (2025)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
by: Kuntz, Thomas, et al.
Published: (2025)
by: Kuntz, Thomas, et al.
Published: (2025)
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
by: Mateega, Spencer, et al.
Published: (2026)
by: Mateega, Spencer, et al.
Published: (2026)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
by: Xia, Chunqiu Steven, et al.
Published: (2025)
by: Xia, Chunqiu Steven, et al.
Published: (2025)
Breaking the Silence: the Threats of Using LLMs in Software Engineering
by: Sallou, June, et al.
Published: (2023)
by: Sallou, June, et al.
Published: (2023)
Can LLMs Replace Manual Annotation of Software Engineering Artifacts?
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
Software Engineering Principles for Fairer Systems: Experiments with GroupCART
by: Peng, Kewen, et al.
Published: (2025)
by: Peng, Kewen, et al.
Published: (2025)
Large Language Models for Software Engineering: A Reproducibility Crisis
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
by: Ding, Yifeng, et al.
Published: (2026)
by: Ding, Yifeng, et al.
Published: (2026)
Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
by: González, Alexandra, et al.
Published: (2025)
by: González, Alexandra, et al.
Published: (2025)
SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
by: Zhao, Zhimin
Published: (2025)
by: Zhao, Zhimin
Published: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)
by: Arora, Avi, et al.
Published: (2025)
Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
by: Golubev, Alexander, et al.
Published: (2025)
by: Golubev, Alexander, et al.
Published: (2025)
Bootstrapping Coding Agents: The Specification Is the Program
by: Monperrus, Martin
Published: (2026)
by: Monperrus, Martin
Published: (2026)
More Rigorous Software Engineering Would Improve Reproducibility in Machine Learning Research
by: Wolter, Moritz, et al.
Published: (2025)
by: Wolter, Moritz, et al.
Published: (2025)
Combating Toxic Language: A Review of LLM-Based Strategies for Software Engineering
by: Zhuo, Hao, et al.
Published: (2025)
by: Zhuo, Hao, et al.
Published: (2025)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026)
by: Yuan, Danlong, et al.
Published: (2026)
The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education
by: Mojica-Hanke, Anamaria, et al.
Published: (2024)
by: Mojica-Hanke, Anamaria, et al.
Published: (2024)
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics
by: Tran, Khang, et al.
Published: (2026)
by: Tran, Khang, et al.
Published: (2026)
Engineering Resource-constrained Software Systems with DNN Components: a Concept-based Pruning Approach
by: Formica, Federico, et al.
Published: (2026)
by: Formica, Federico, et al.
Published: (2026)
Foundation Model Engineering: Engineering Foundation Models Just as Engineering Software
by: Ran, Dezhi, et al.
Published: (2024)
by: Ran, Dezhi, et al.
Published: (2024)
On Problems of Implicit Context Compression for Software Engineering Agents
by: Gelvan, Kirill, et al.
Published: (2026)
by: Gelvan, Kirill, et al.
Published: (2026)
Agentless: Demystifying LLM-based Software Engineering Agents
by: Xia, Chunqiu Steven, et al.
Published: (2024)
by: Xia, Chunqiu Steven, et al.
Published: (2024)
Can Coding Agents Be General Agents?
by: Ivanov, Maksim, et al.
Published: (2026)
by: Ivanov, Maksim, et al.
Published: (2026)
A Machine Learning-Based Error Mitigation Approach For Reliable Software Development On IBM'S Quantum Computers
by: Muqeet, Asmar, et al.
Published: (2024)
by: Muqeet, Asmar, et al.
Published: (2024)
Should Code Models Learn Pedagogically? A Preliminary Evaluation of Curriculum Learning for Real-World Software Engineering Tasks
by: Khant, Kyi Shin, et al.
Published: (2025)
by: Khant, Kyi Shin, et al.
Published: (2025)
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
by: Zhang, Kexun, et al.
Published: (2024)
by: Zhang, Kexun, et al.
Published: (2024)
Challenges and Paths Towards AI for Software Engineering
by: Gu, Alex, et al.
Published: (2025)
by: Gu, Alex, et al.
Published: (2025)
On the Replicability and Reproducibility of Deep Learning in Software Engineering
by: Liu, Chao, et al.
Published: (2020)
by: Liu, Chao, et al.
Published: (2020)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
by: Yang, John, et al.
Published: (2024)
by: Yang, John, et al.
Published: (2024)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
by: LeVine, Will, et al.
Published: (2026)
by: LeVine, Will, et al.
Published: (2026)
Similar Items
-
miniCodeProps: a Minimal Benchmark for Proving Code Properties
by: Lohn, Evan, et al.
Published: (2024) -
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026) -
Assessing the Use of AutoML for Data-Driven Software Engineering
by: Calefato, Fabio, et al.
Published: (2023) -
Agint: Agentic Graph Compilation for Software Engineering Agents
by: Chivukula, Abhi, et al.
Published: (2025) -
A Systematic Literature Review on the Use of Machine Learning in Software Engineering
by: Fred, Nyaga, et al.
Published: (2024)