From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
Fuente:
arXiv
Saved in:
| Main Authors: | Rajbahadur, Gopi Krishnan, Oliva, Gustavo A., Lin, Dayi, Shin, Jiho, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
by: Vasilevski, Kirill, et al.
Published: (2025)
by: Vasilevski, Kirill, et al.
Published: (2025)
Data Quality Antipatterns for Software Analytics
by: Bhatia, Aaditya, et al.
Published: (2024)
by: Bhatia, Aaditya, et al.
Published: (2024)
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025)
by: Jewitt, James, et al.
Published: (2025)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
by: Oliva, Gustavo A., et al.
Published: (2025)
by: Oliva, Gustavo A., et al.
Published: (2025)
Keeping Deep Learning Models in Check: A History-Based Approach to Mitigate Overfitting
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
by: Ebrahimi, Amir M., et al.
Published: (2026)
by: Ebrahimi, Amir M., et al.
Published: (2026)
SimClone: Detecting Tabular Data Clones using Value Similarity
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026)
by: Jewitt, James, et al.
Published: (2026)
LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance
by: Tan, Jingwen, et al.
Published: (2024)
by: Tan, Jingwen, et al.
Published: (2024)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
by: Hassan, Ahmed E., et al.
Published: (2025)
by: Hassan, Ahmed E., et al.
Published: (2025)
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
by: Chen, Zhilong, et al.
Published: (2025)
by: Chen, Zhilong, et al.
Published: (2025)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
Engineering AI Judge Systems
by: Lin, Jiahuei, et al.
Published: (2024)
by: Lin, Jiahuei, et al.
Published: (2024)
Building an Open AIBOM Standard in the Wild
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2025)
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2025)
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement
by: Gallaba, Keheliya, et al.
Published: (2025)
by: Gallaba, Keheliya, et al.
Published: (2025)
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
by: Rombaut, Benjamin, et al.
Published: (2024)
by: Rombaut, Benjamin, et al.
Published: (2024)
What Slows Down FMware Development? An Empirical Study of Developer Challenges and Resolution Times
by: Wang, Zitao, et al.
Published: (2025)
by: Wang, Zitao, et al.
Published: (2025)
Implementing AI Bill of Materials (AI BOM) with SPDX 3.0: A Comprehensive Guide to Creating AI and Dataset Bill of Materials
by: Bennet, Karen, et al.
Published: (2025)
by: Bennet, Karen, et al.
Published: (2025)
When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models
by: Zheng, Shenyu, et al.
Published: (2026)
by: Zheng, Shenyu, et al.
Published: (2026)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
by: Fan, Zhiyu, et al.
Published: (2025)
by: Fan, Zhiyu, et al.
Published: (2025)
Domain Adaptation for Code Model-based Unit Test Case Generation
by: Shin, Jiho, et al.
Published: (2023)
by: Shin, Jiho, et al.
Published: (2023)
Predicting long time contributors with knowledge units of programming languages: an empirical study
by: Ahasanuzzaman, Md, et al.
Published: (2024)
by: Ahasanuzzaman, Md, et al.
Published: (2024)
Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap
by: Sartaj, Hassan, et al.
Published: (2025)
by: Sartaj, Hassan, et al.
Published: (2025)
Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
by: Cogo, Filipe R., et al.
Published: (2025)
by: Cogo, Filipe R., et al.
Published: (2025)
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
Retrieval-Augmented Test Generation: How Far Are We?
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
Towards Reliable Generation of Executable Workflows by Foundation Models
by: Masoumzadeh, Sogol, et al.
Published: (2025)
by: Masoumzadeh, Sogol, et al.
Published: (2025)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
by: Bian, Yutong, et al.
Published: (2025)
by: Bian, Yutong, et al.
Published: (2025)
Predicting post-release defects with knowledge units (KUs) of programming languages: an empirical study
by: Ahasanuzzaman, Md, et al.
Published: (2024)
by: Ahasanuzzaman, Md, et al.
Published: (2024)
Causal Software Engineering: A Vision and Roadmap
by: Pietrantuono, Roberto, et al.
Published: (2026)
by: Pietrantuono, Roberto, et al.
Published: (2026)
AppForge: From Assistant to Independent Developer -- Are GPTs Ready for Software Development?
by: Ran, Dezhi, et al.
Published: (2025)
by: Ran, Dezhi, et al.
Published: (2025)
AIDev: Studying AI Coding Agents on GitHub
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study
by: Ahasanuzzaman, Md, et al.
Published: (2026)
by: Ahasanuzzaman, Md, et al.
Published: (2026)
Similar Items
-
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
by: Hassan, Ahmed E., et al.
Published: (2024) -
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
by: Vasilevski, Kirill, et al.
Published: (2025) -
Data Quality Antipatterns for Software Analytics
by: Bhatia, Aaditya, et al.
Published: (2024) -
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
by: Hassan, Ahmed E., et al.
Published: (2024) -
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025)