Data Quality Cascades in AI: Characterizing, Mitigating, and Preventing Error Propagation
Fuente:
Zenodo
Enregistré dans:
| Auteurs principaux: | , |
|---|---|
| Format: | Recurso digital |
| Publié: |
Zenodo
2025
|
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866901165686390784 |
|---|---|
| author | Revista, Zen IA, 10 |
| author_facet | Revista, Zen IA, 10 |
| contents | This paper explores the phenomenon of data quality cascades in artificial intelligence (AI) systems, focusing on the characterization, mitigation, and prevention of error propagation. Data quality is paramount to the performance and reliability of AI models. However, errors and inconsistencies in data can propagate through various stages of the AI pipeline, creating a cascade effect that significantly degrades model accuracy and decision-making capabilities. We investigate the sources and types of data quality issues that contribute to these cascades, including biases, noise, incompleteness, and inconsistencies. Furthermore, we examine the mechanisms through which these errors propagate, such as feature engineering, model training, and deployment. We propose strategies for mitigating the impact of data quality cascades, including data cleaning techniques, bias detection and correction methods, robust model training approaches, and error monitoring systems. Finally, we discuss preventative measures aimed at ensuring high data quality throughout the AI lifecycle, such as data governance frameworks, automated data validation procedures, and data provenance tracking. Through a combination of theoretical analysis and practical examples, this paper provides a comprehensive understanding of data quality cascades in AI and offers actionable insights for building more robust and reliable AI systems. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17829843 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Data Quality Cascades in AI: Characterizing, Mitigating, and Preventing Error Propagation Revista, Zen IA, 10 This paper explores the phenomenon of data quality cascades in artificial intelligence (AI) systems, focusing on the characterization, mitigation, and prevention of error propagation. Data quality is paramount to the performance and reliability of AI models. However, errors and inconsistencies in data can propagate through various stages of the AI pipeline, creating a cascade effect that significantly degrades model accuracy and decision-making capabilities. We investigate the sources and types of data quality issues that contribute to these cascades, including biases, noise, incompleteness, and inconsistencies. Furthermore, we examine the mechanisms through which these errors propagate, such as feature engineering, model training, and deployment. We propose strategies for mitigating the impact of data quality cascades, including data cleaning techniques, bias detection and correction methods, robust model training approaches, and error monitoring systems. Finally, we discuss preventative measures aimed at ensuring high data quality throughout the AI lifecycle, such as data governance frameworks, automated data validation procedures, and data provenance tracking. Through a combination of theoretical analysis and practical examples, this paper provides a comprehensive understanding of data quality cascades in AI and offers actionable insights for building more robust and reliable AI systems. |
| title | Data Quality Cascades in AI: Characterizing, Mitigating, and Preventing Error Propagation |
| url | https://doi.org/10.5281/zenodo.17829843 |