Data Quality Antipatterns for Software Analytics
Fuente:
arXiv
Saved in:
| Main Authors: | Bhatia, Aaditya, Lin, Dayi, Rajbahadur, Gopi Krishnan, Adams, Bram, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025)
by: Jewitt, James, et al.
Published: (2025)
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
by: Oliva, Gustavo A., et al.
Published: (2025)
by: Oliva, Gustavo A., et al.
Published: (2025)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026)
by: Jewitt, James, et al.
Published: (2026)
Keeping Deep Learning Models in Check: A History-Based Approach to Mitigate Overfitting
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
by: Ebrahimi, Amir M., et al.
Published: (2026)
by: Ebrahimi, Amir M., et al.
Published: (2026)
SimClone: Detecting Tabular Data Clones using Value Similarity
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
by: Hassan, Ahmed E., et al.
Published: (2025)
by: Hassan, Ahmed E., et al.
Published: (2025)
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
An Empirical Study of Self-Admitted Technical Debt in Machine Learning Software
by: Bhatia, Aaditya, et al.
Published: (2023)
by: Bhatia, Aaditya, et al.
Published: (2023)
LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance
by: Tan, Jingwen, et al.
Published: (2024)
by: Tan, Jingwen, et al.
Published: (2024)
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
by: Vasilevski, Kirill, et al.
Published: (2025)
by: Vasilevski, Kirill, et al.
Published: (2025)
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
by: Hasan, Mohammed Mehedi, et al.
Published: (2026)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Engineering AI Judge Systems
by: Lin, Jiahuei, et al.
Published: (2024)
by: Lin, Jiahuei, et al.
Published: (2024)
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
by: Hasan, Mohammed Mehedi, et al.
Published: (2025)
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
by: Chen, Zhilong, et al.
Published: (2025)
by: Chen, Zhilong, et al.
Published: (2025)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
OmniLLP: Enhancing LLM-based Log Level Prediction with Context-Aware Retrieval
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
by: Rombaut, Benjamin, et al.
Published: (2024)
by: Rombaut, Benjamin, et al.
Published: (2024)
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement
by: Gallaba, Keheliya, et al.
Published: (2025)
by: Gallaba, Keheliya, et al.
Published: (2025)
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
by: Fan, Zhiyu, et al.
Published: (2025)
by: Fan, Zhiyu, et al.
Published: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026)
by: Zhao, Zhimin, et al.
Published: (2026)
Software Performance Engineering for Foundation Model-Powered Software
by: Zhang, Haoxiang, et al.
Published: (2024)
by: Zhang, Haoxiang, et al.
Published: (2024)
Building an Open AIBOM Standard in the Wild
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2025)
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2025)
Software Engineering and Foundation Models: Insights from Industry Blogs Using a Jury of Foundation Models
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Output Format Biases in the Evaluation of Large Language Models for Code Translation
by: Macedo, Marcos, et al.
Published: (2024)
by: Macedo, Marcos, et al.
Published: (2024)
Implementing AI Bill of Materials (AI BOM) with SPDX 3.0: A Comprehensive Guide to Creating AI and Dataset Bill of Materials
by: Bennet, Karen, et al.
Published: (2025)
by: Bennet, Karen, et al.
Published: (2025)
InterTrans: Leveraging Transitive Intermediate Translations to Enhance LLM-based Code Translation
by: Macedo, Marcos, et al.
Published: (2024)
by: Macedo, Marcos, et al.
Published: (2024)
The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Does Using Bazel Help Speed Up Continuous Integration Builds?
by: Zheng, Shenyu, et al.
Published: (2024)
by: Zheng, Shenyu, et al.
Published: (2024)
AIDev: Studying AI Coding Agents on GitHub
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Future of Artificial Intelligence in Agile Software Development
by: Mahboob, Mariyam, et al.
Published: (2024)
by: Mahboob, Mariyam, et al.
Published: (2024)
Similar Items
-
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025) -
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024) -
Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
by: Li, Hao, et al.
Published: (2024) -
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
by: Oliva, Gustavo A., et al.
Published: (2025) -
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026)