Salvato in:
| Autori principali: | He, Kang, Roy, Kaushik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2603.01327 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SWE-Exp: Experience-Driven Software Issue Resolution
di: Chen, Silin, et al.
Pubblicazione: (2025)
di: Chen, Silin, et al.
Pubblicazione: (2025)
SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
di: Li, Han, et al.
Pubblicazione: (2025)
di: Li, Han, et al.
Pubblicazione: (2025)
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
di: Wong, Sherman, et al.
Pubblicazione: (2025)
di: Wong, Sherman, et al.
Pubblicazione: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
di: Raghavendra, Mohit, et al.
Pubblicazione: (2026)
di: Raghavendra, Mohit, et al.
Pubblicazione: (2026)
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
di: Chen, Jialong, et al.
Pubblicazione: (2026)
di: Chen, Jialong, et al.
Pubblicazione: (2026)
HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench
di: Wang, Yueyang, et al.
Pubblicazione: (2026)
di: Wang, Yueyang, et al.
Pubblicazione: (2026)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
di: Jain, Naman, et al.
Pubblicazione: (2025)
di: Jain, Naman, et al.
Pubblicazione: (2025)
Exploring LLM-based Agents for Root Cause Analysis
di: Roy, Devjeet, et al.
Pubblicazione: (2024)
di: Roy, Devjeet, et al.
Pubblicazione: (2024)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
di: Wang, Lilin, et al.
Pubblicazione: (2025)
di: Wang, Lilin, et al.
Pubblicazione: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
di: Shetty, Manish, et al.
Pubblicazione: (2025)
di: Shetty, Manish, et al.
Pubblicazione: (2025)
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
di: Peng, Jinjun, et al.
Pubblicazione: (2026)
di: Peng, Jinjun, et al.
Pubblicazione: (2026)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
di: Wei, Yuxiang, et al.
Pubblicazione: (2025)
di: Wei, Yuxiang, et al.
Pubblicazione: (2025)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
di: Xia, Chunqiu Steven, et al.
Pubblicazione: (2025)
di: Xia, Chunqiu Steven, et al.
Pubblicazione: (2025)
FormulaCode: Evaluating Agentic Optimization on Large Codebases
di: Sehgal, Atharva, et al.
Pubblicazione: (2026)
di: Sehgal, Atharva, et al.
Pubblicazione: (2026)
SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents
di: Kon, Patrick Tser Jern, et al.
Pubblicazione: (2026)
di: Kon, Patrick Tser Jern, et al.
Pubblicazione: (2026)
Aligning the Objective of LLM-based Program Repair
di: Xu, Junjielong, et al.
Pubblicazione: (2024)
di: Xu, Junjielong, et al.
Pubblicazione: (2024)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
di: Yang, John, et al.
Pubblicazione: (2024)
di: Yang, John, et al.
Pubblicazione: (2024)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
di: Hassid, Michael, et al.
Pubblicazione: (2024)
di: Hassid, Michael, et al.
Pubblicazione: (2024)
New Solutions on LLM Acceleration, Optimization, and Application
di: Huang, Yingbing, et al.
Pubblicazione: (2024)
di: Huang, Yingbing, et al.
Pubblicazione: (2024)
CP-Agent: Agentic Constraint Programming
di: Szeider, Stefan
Pubblicazione: (2025)
di: Szeider, Stefan
Pubblicazione: (2025)
AFlow: Automating Agentic Workflow Generation
di: Zhang, Jiayi, et al.
Pubblicazione: (2024)
di: Zhang, Jiayi, et al.
Pubblicazione: (2024)
Wisdom and Delusion of LLM Ensembles for Code Generation and Repair
di: Vallecillos-Ruiz, Fernando, et al.
Pubblicazione: (2025)
di: Vallecillos-Ruiz, Fernando, et al.
Pubblicazione: (2025)
SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
di: Badertdinov, Ibragim, et al.
Pubblicazione: (2026)
di: Badertdinov, Ibragim, et al.
Pubblicazione: (2026)
OptLLM: Optimal Assignment of Queries to Large Language Models
di: Liu, Yueyue, et al.
Pubblicazione: (2024)
di: Liu, Yueyue, et al.
Pubblicazione: (2024)
Scaling Test-Time Compute for Agentic Coding
di: Kim, Joongwon, et al.
Pubblicazione: (2026)
di: Kim, Joongwon, et al.
Pubblicazione: (2026)
Let the Code LLM Edit Itself When You Edit the Code
di: He, Zhenyu, et al.
Pubblicazione: (2024)
di: He, Zhenyu, et al.
Pubblicazione: (2024)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
di: Peng, Jinjun, et al.
Pubblicazione: (2025)
di: Peng, Jinjun, et al.
Pubblicazione: (2025)
SWE-bench Goes Live!
di: Zhang, Linghao, et al.
Pubblicazione: (2025)
di: Zhang, Linghao, et al.
Pubblicazione: (2025)
Otter: Generating Tests from Issues to Validate SWE Patches
di: Ahmed, Toufique, et al.
Pubblicazione: (2025)
di: Ahmed, Toufique, et al.
Pubblicazione: (2025)
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
di: Wu, Jie JW, et al.
Pubblicazione: (2025)
di: Wu, Jie JW, et al.
Pubblicazione: (2025)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
di: Martinez, Matias
Pubblicazione: (2024)
di: Martinez, Matias
Pubblicazione: (2024)
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
di: Weyssow, Martin, et al.
Pubblicazione: (2024)
di: Weyssow, Martin, et al.
Pubblicazione: (2024)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
di: Guo, Daya, et al.
Pubblicazione: (2024)
di: Guo, Daya, et al.
Pubblicazione: (2024)
Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence
di: Song, Linxin, et al.
Pubblicazione: (2026)
di: Song, Linxin, et al.
Pubblicazione: (2026)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
di: Ahmed, Toufique, et al.
Pubblicazione: (2024)
di: Ahmed, Toufique, et al.
Pubblicazione: (2024)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
di: Jiang, Juyong, et al.
Pubblicazione: (2026)
di: Jiang, Juyong, et al.
Pubblicazione: (2026)
Refining GPT-3 Embeddings with a Siamese Structure for Technical Post Duplicate Detection
di: Wu, Xingfang, et al.
Pubblicazione: (2023)
di: Wu, Xingfang, et al.
Pubblicazione: (2023)
SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving
di: Tao, Chaofan, et al.
Pubblicazione: (2026)
di: Tao, Chaofan, et al.
Pubblicazione: (2026)
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
di: Wang, Shuaiqi, et al.
Pubblicazione: (2026)
di: Wang, Shuaiqi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SWE-Exp: Experience-Driven Software Issue Resolution
di: Chen, Silin, et al.
Pubblicazione: (2025) -
SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
di: Li, Han, et al.
Pubblicazione: (2025) -
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
di: Wong, Sherman, et al.
Pubblicazione: (2025) -
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
di: Raghavendra, Mohit, et al.
Pubblicazione: (2026) -
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
di: Chen, Jialong, et al.
Pubblicazione: (2026)