SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Zhao, Zhimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
von: Yuan, Danlong, et al.
Veröffentlicht: (2026)
von: Yuan, Danlong, et al.
Veröffentlicht: (2026)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
von: Ding, Yifeng, et al.
Veröffentlicht: (2026)
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
von: Miserendino, Samuel, et al.
Veröffentlicht: (2025)
von: Miserendino, Samuel, et al.
Veröffentlicht: (2025)
Foundation Model Engineering: Engineering Foundation Models Just as Engineering Software
von: Ran, Dezhi, et al.
Veröffentlicht: (2024)
von: Ran, Dezhi, et al.
Veröffentlicht: (2024)
SWE-Exp: Experience-Driven Software Issue Resolution
von: Chen, Silin, et al.
Veröffentlicht: (2025)
von: Chen, Silin, et al.
Veröffentlicht: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2025)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2025)
Investigating Test Overfitting on SWE-bench
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
von: Li, Han, et al.
Veröffentlicht: (2025)
von: Li, Han, et al.
Veröffentlicht: (2025)
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
von: Zhao, Zhimin, et al.
Veröffentlicht: (2024)
von: Zhao, Zhimin, et al.
Veröffentlicht: (2024)
SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents
von: Kon, Patrick Tser Jern, et al.
Veröffentlicht: (2026)
von: Kon, Patrick Tser Jern, et al.
Veröffentlicht: (2026)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
von: Zhao, Zhimin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhimin, et al.
Veröffentlicht: (2026)
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
von: Mateega, Spencer, et al.
Veröffentlicht: (2026)
von: Mateega, Spencer, et al.
Veröffentlicht: (2026)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
Otter: Generating Tests from Issues to Validate SWE Patches
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2026)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2026)
Large Language Models for Software Engineering: A Reproducibility Crisis
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2025)
von: Siddiq, Mohammed Latif, et al.
Veröffentlicht: (2025)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
Should Code Models Learn Pedagogically? A Preliminary Evaluation of Curriculum Learning for Real-World Software Engineering Tasks
von: Khant, Kyi Shin, et al.
Veröffentlicht: (2025)
von: Khant, Kyi Shin, et al.
Veröffentlicht: (2025)
Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
von: González, Alexandra, et al.
Veröffentlicht: (2025)
von: González, Alexandra, et al.
Veröffentlicht: (2025)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2026)
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2026)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
Agint: Agentic Graph Compilation for Software Engineering Agents
von: Chivukula, Abhi, et al.
Veröffentlicht: (2025)
von: Chivukula, Abhi, et al.
Veröffentlicht: (2025)
Breaking the Silence: the Threats of Using LLMs in Software Engineering
von: Sallou, June, et al.
Veröffentlicht: (2023)
von: Sallou, June, et al.
Veröffentlicht: (2023)
Software Engineering Principles for Fairer Systems: Experiments with GroupCART
von: Peng, Kewen, et al.
Veröffentlicht: (2025)
von: Peng, Kewen, et al.
Veröffentlicht: (2025)
Assessing the Use of AutoML for Data-Driven Software Engineering
von: Calefato, Fabio, et al.
Veröffentlicht: (2023)
von: Calefato, Fabio, et al.
Veröffentlicht: (2023)
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
von: Bula, Timothy, et al.
Veröffentlicht: (2025)
von: Bula, Timothy, et al.
Veröffentlicht: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
A Systematic Literature Review on the Use of Machine Learning in Software Engineering
von: Fred, Nyaga, et al.
Veröffentlicht: (2024)
von: Fred, Nyaga, et al.
Veröffentlicht: (2024)
SWE-Bench-CL: Continual Learning for Coding Agents
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
Sentiment Analysis in Software Engineering: Evaluating Generative Pre-trained Transformers
von: Saifullah, KM Khalid, et al.
Veröffentlicht: (2025)
von: Saifullah, KM Khalid, et al.
Veröffentlicht: (2025)
More Rigorous Software Engineering Would Improve Reproducibility in Machine Learning Research
von: Wolter, Moritz, et al.
Veröffentlicht: (2025)
von: Wolter, Moritz, et al.
Veröffentlicht: (2025)
Combating Toxic Language: A Review of LLM-Based Strategies for Software Engineering
von: Zhuo, Hao, et al.
Veröffentlicht: (2025)
von: Zhuo, Hao, et al.
Veröffentlicht: (2025)
Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education
von: Mojica-Hanke, Anamaria, et al.
Veröffentlicht: (2024)
von: Mojica-Hanke, Anamaria, et al.
Veröffentlicht: (2024)
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
Engineering Resource-constrained Software Systems with DNN Components: a Concept-based Pruning Approach
von: Formica, Federico, et al.
Veröffentlicht: (2026)
von: Formica, Federico, et al.
Veröffentlicht: (2026)
Toward Explaining Large Language Models in Software Engineering Tasks
von: Vitale, Antonio, et al.
Veröffentlicht: (2025)
von: Vitale, Antonio, et al.
Veröffentlicht: (2025)
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
von: Vasilevski, Kirill, et al.
Veröffentlicht: (2025)
von: Vasilevski, Kirill, et al.
Veröffentlicht: (2025)
A Model-Driven Engineering Approach to AI-Powered Healthcare Platforms
von: Raheem, Mira, et al.
Veröffentlicht: (2025)
von: Raheem, Mira, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
von: Yuan, Danlong, et al.
Veröffentlicht: (2026) -
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
von: Ding, Yifeng, et al.
Veröffentlicht: (2026) -
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
von: Miserendino, Samuel, et al.
Veröffentlicht: (2025) -
Foundation Model Engineering: Engineering Foundation Models Just as Engineering Software
von: Ran, Dezhi, et al.
Veröffentlicht: (2024) -
SWE-Exp: Experience-Driven Software Issue Resolution
von: Chen, Silin, et al.
Veröffentlicht: (2025)