Gespeichert in:
| Hauptverfasser: | Oliva, Gustavo A., Rajbahadur, Gopi Krishnan, Bhatia, Aaditya, Zhang, Haoxiang, Chen, Yihao, Chen, Zhilong, Leung, Arthur, Lin, Dayi, Chen, Boyuan, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.09108 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
von: Chen, Zhilong, et al.
Veröffentlicht: (2025)
von: Chen, Zhilong, et al.
Veröffentlicht: (2025)
Data Quality Antipatterns for Software Analytics
von: Bhatia, Aaditya, et al.
Veröffentlicht: (2024)
von: Bhatia, Aaditya, et al.
Veröffentlicht: (2024)
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
von: Rajbahadur, Gopi Krishnan, et al.
Veröffentlicht: (2024)
von: Rajbahadur, Gopi Krishnan, et al.
Veröffentlicht: (2024)
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
von: Hassan, Ahmed E., et al.
Veröffentlicht: (2024)
von: Hassan, Ahmed E., et al.
Veröffentlicht: (2024)
SimClone: Detecting Tabular Data Clones using Value Similarity
von: Yang, Xu, et al.
Veröffentlicht: (2024)
von: Yang, Xu, et al.
Veröffentlicht: (2024)
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
von: Ebrahimi, Amir M., et al.
Veröffentlicht: (2026)
von: Ebrahimi, Amir M., et al.
Veröffentlicht: (2026)
Keeping Deep Learning Models in Check: A History-Based Approach to Mitigate Overfitting
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
PromptExp: Multi-granularity Prompt Explanation of Large Language Models
von: Dong, Ximing, et al.
Veröffentlicht: (2024)
von: Dong, Ximing, et al.
Veröffentlicht: (2024)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
von: Fan, Zhiyu, et al.
Veröffentlicht: (2025)
von: Fan, Zhiyu, et al.
Veröffentlicht: (2025)
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
von: Vasilevski, Kirill, et al.
Veröffentlicht: (2025)
von: Vasilevski, Kirill, et al.
Veröffentlicht: (2025)
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
von: Hassan, Ahmed E., et al.
Veröffentlicht: (2024)
von: Hassan, Ahmed E., et al.
Veröffentlicht: (2024)
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
von: Hassan, Ahmed E., et al.
Veröffentlicht: (2024)
von: Hassan, Ahmed E., et al.
Veröffentlicht: (2024)
Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
von: Jewitt, James, et al.
Veröffentlicht: (2025)
von: Jewitt, James, et al.
Veröffentlicht: (2025)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
von: Jewitt, James, et al.
Veröffentlicht: (2026)
von: Jewitt, James, et al.
Veröffentlicht: (2026)
Implementing AI Bill of Materials (AI BOM) with SPDX 3.0: A Comprehensive Guide to Creating AI and Dataset Bill of Materials
von: Bennet, Karen, et al.
Veröffentlicht: (2025)
von: Bennet, Karen, et al.
Veröffentlicht: (2025)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2026)
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2026)
Software Performance Engineering for Foundation Model-Powered Software
von: Zhang, Haoxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Haoxiang, et al.
Veröffentlicht: (2024)
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2025)
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2025)
An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2025)
von: Hasan, Mohammed Mehedi, et al.
Veröffentlicht: (2025)
SLA-Awareness for AI-assisted coding
von: Thangarajah, Kishanthan, et al.
Veröffentlicht: (2025)
von: Thangarajah, Kishanthan, et al.
Veröffentlicht: (2025)
Building an Open AIBOM Standard in the Wild
von: Rajbahadur, Gopi Krishnan, et al.
Veröffentlicht: (2025)
von: Rajbahadur, Gopi Krishnan, et al.
Veröffentlicht: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
von: Dong, Ximing, et al.
Veröffentlicht: (2026)
von: Dong, Ximing, et al.
Veröffentlicht: (2026)
When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models
von: Zheng, Shenyu, et al.
Veröffentlicht: (2026)
von: Zheng, Shenyu, et al.
Veröffentlicht: (2026)
LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance
von: Tan, Jingwen, et al.
Veröffentlicht: (2024)
von: Tan, Jingwen, et al.
Veröffentlicht: (2024)
What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair
von: Martinez, Matias, et al.
Veröffentlicht: (2026)
von: Martinez, Matias, et al.
Veröffentlicht: (2026)
An Empirical Study of Self-Admitted Technical Debt in Machine Learning Software
von: Bhatia, Aaditya, et al.
Veröffentlicht: (2023)
von: Bhatia, Aaditya, et al.
Veröffentlicht: (2023)
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2025)
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2025)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
Virasoro Blocks and Trouble at the Euclidean Horizon
von: Datar, Aaditya, et al.
Veröffentlicht: (2025)
von: Datar, Aaditya, et al.
Veröffentlicht: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks
von: Chen, Jiao, et al.
Veröffentlicht: (2026)
von: Chen, Jiao, et al.
Veröffentlicht: (2026)
Active Learning to Guide Labeling Efforts for Question Difficulty Estimation
von: Thuy, Arthur, et al.
Veröffentlicht: (2024)
von: Thuy, Arthur, et al.
Veröffentlicht: (2024)
Multi-Agent Clarity-Aware Dynamic Coverage with Gaussian Processes
von: Agrawal, Devansh R., et al.
Veröffentlicht: (2024)
von: Agrawal, Devansh R., et al.
Veröffentlicht: (2024)
A Review Paper on Use of Recycled Aggregates in Concrete
von: Neeraj Kumar, et al.
Veröffentlicht: (2022)
von: Neeraj Kumar, et al.
Veröffentlicht: (2022)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
von: Deng, Xiang, et al.
Veröffentlicht: (2025)
von: Deng, Xiang, et al.
Veröffentlicht: (2025)
Clarity and Computational Efficiency of Orbital Boundary Labeling
von: Wallinger, Markus, et al.
Veröffentlicht: (2026)
von: Wallinger, Markus, et al.
Veröffentlicht: (2026)
Gradient dynamics for low-rank fine-tuning beyond kernels
von: Dayi, Arif Kerem, et al.
Veröffentlicht: (2024)
von: Dayi, Arif Kerem, et al.
Veröffentlicht: (2024)
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
von: Xu, Tian, et al.
Veröffentlicht: (2024)
von: Xu, Tian, et al.
Veröffentlicht: (2024)
Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms
von: Xu, Tian, et al.
Veröffentlicht: (2026)
von: Xu, Tian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
von: Chen, Zhilong, et al.
Veröffentlicht: (2025) -
Data Quality Antipatterns for Software Analytics
von: Bhatia, Aaditya, et al.
Veröffentlicht: (2024) -
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
von: Rajbahadur, Gopi Krishnan, et al.
Veröffentlicht: (2024) -
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
von: Hassan, Ahmed E., et al.
Veröffentlicht: (2024) -
SimClone: Detecting Tabular Data Clones using Value Similarity
von: Yang, Xu, et al.
Veröffentlicht: (2024)