Limitations of Public Chest Radiography Datasets for Artificial Intelligence: Label Quality, Domain Shift, Bias and Evaluation Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Rafferty, Amy, Rajan, Ajitha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Expert Input for Robust and Explainable AI-Assisted Lung Cancer Detection in Chest X-rays
by: Rafferty, Amy, et al.
Published: (2024)
by: Rafferty, Amy, et al.
Published: (2024)
Simbanex: Similarity-based Exploration of IEEE VIS Publications
by: Witschard, Daniel, et al.
Published: (2024)
by: Witschard, Daniel, et al.
Published: (2024)
FedAIoT: A Federated Learning Benchmark for Artificial Intelligence of Things
by: Alam, Samiul, et al.
Published: (2023)
by: Alam, Samiul, et al.
Published: (2023)
Sanidha: A Studio Quality Multi-Modal Dataset for Carnatic Music
by: Krishnan, Venkatakrishnan Vaidyanathapuram, et al.
Published: (2025)
by: Krishnan, Venkatakrishnan Vaidyanathapuram, et al.
Published: (2025)
ARED: Argentina Real Estate Dataset
by: Belenky, Iván
Published: (2024)
by: Belenky, Iván
Published: (2024)
A Gold Standard Dataset for the Reviewer Assignment Problem
by: Stelmakh, Ivan, et al.
Published: (2023)
by: Stelmakh, Ivan, et al.
Published: (2023)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
by: Rao, Delip, et al.
Published: (2024)
by: Rao, Delip, et al.
Published: (2024)
Using Artificial Intuition in Distinct, Minimalist Classification of Scientific Abstracts for Management of Technology Portfolios
by: Ranka, Prateek, et al.
Published: (2025)
by: Ranka, Prateek, et al.
Published: (2025)
BatteryLife: A Comprehensive Dataset and Benchmark for Battery Life Prediction
by: Tan, Ruifeng, et al.
Published: (2025)
by: Tan, Ruifeng, et al.
Published: (2025)
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
by: Cheng, Ming, et al.
Published: (2024)
by: Cheng, Ming, et al.
Published: (2024)
Quantum Machine Learning: Unveiling Trends, Impacts through Bibliometric Analysis
by: Bansal, Riya, et al.
Published: (2025)
by: Bansal, Riya, et al.
Published: (2025)
A Causal Inference Approach for Quantifying Research Impact
by: Ochiai, Keiichi, et al.
Published: (2025)
by: Ochiai, Keiichi, et al.
Published: (2025)
Clustering scientific publications: lessons learned through experiments with a real citation network
by: Huong, Vu Thi, et al.
Published: (2025)
by: Huong, Vu Thi, et al.
Published: (2025)
Gymnasium: A Standard Interface for Reinforcement Learning Environments
by: Towers, Mark, et al.
Published: (2024)
by: Towers, Mark, et al.
Published: (2024)
ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review
by: Riehl, Kevin, et al.
Published: (2026)
by: Riehl, Kevin, et al.
Published: (2026)
The IJCNN 2025 Review Process
by: Scarpiniti, Michele, et al.
Published: (2026)
by: Scarpiniti, Michele, et al.
Published: (2026)
Beyond coauthorship: semantic structure and phantom collaborators in transportation research, 1967--2025
by: Choi, Seongjin
Published: (2026)
by: Choi, Seongjin
Published: (2026)
Benchmark Data Repositories for Better Benchmarking
by: Longjohn, Rachel, et al.
Published: (2024)
by: Longjohn, Rachel, et al.
Published: (2024)
Author Name Disambiguation via Heterogeneous Network Embedding from Structural and Semantic Perspectives
by: Xie, Wenjin, et al.
Published: (2022)
by: Xie, Wenjin, et al.
Published: (2022)
Quantitative Methods in Research Evaluation Citation Indicators, Altmetrics, and Artificial Intelligence
by: Thelwall, Mike
Published: (2024)
by: Thelwall, Mike
Published: (2024)
A Bibliometric Analysis of Highly Cited Artificial Intelligence Publications in Science Citation Index Expanded
by: Ho, Yuh-Shan, et al.
Published: (2024)
by: Ho, Yuh-Shan, et al.
Published: (2024)
Radiologist-Guided Causal Concept Bottleneck Models for Chest X-Ray Interpretation
by: Rafferty, Amy, et al.
Published: (2026)
by: Rafferty, Amy, et al.
Published: (2026)
Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets
by: Attrach, Rafi Al, et al.
Published: (2026)
by: Attrach, Rafi Al, et al.
Published: (2026)
Clinically Aware Synthetic Image Generation for Concept Coverage in Chest X-ray Models
by: Rafferty, Amy, et al.
Published: (2026)
by: Rafferty, Amy, et al.
Published: (2026)
CoRPA: Adversarial Image Generation for Chest X-rays Using Concept Vector Perturbations and Generative Models
by: Rafferty, Amy, et al.
Published: (2025)
by: Rafferty, Amy, et al.
Published: (2025)
OpenDataLab: Empowering General Artificial Intelligence with Open Datasets
by: He, Conghui, et al.
Published: (2024)
by: He, Conghui, et al.
Published: (2024)
Publication Trends in Artificial Intelligence Conferences: The Rise of Super Prolific Authors
by: Azad, Ariful, et al.
Published: (2024)
by: Azad, Ariful, et al.
Published: (2024)
Rejoinder: The ICML 2023 Ranking Experiment: Examining Author Self-Assessment in ML/AI Peer Review
by: Su, Buxin, et al.
Published: (2026)
by: Su, Buxin, et al.
Published: (2026)
Labeling Case Similarity based on Co-Citation of Legal Articles in Judgment Documents with Empirical Dispute-Based Evaluation
by: Liu, Chao-Lin, et al.
Published: (2025)
by: Liu, Chao-Lin, et al.
Published: (2025)
Dissecting Submission Limit in Desk-Rejections: A Mathematical Analysis of Fairness in AI Conference Policies
by: Cao, Yuefan, et al.
Published: (2025)
by: Cao, Yuefan, et al.
Published: (2025)
Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias
by: Algaba, Andres, et al.
Published: (2024)
by: Algaba, Andres, et al.
Published: (2024)
Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
by: Smetana, Mason, et al.
Published: (2025)
by: Smetana, Mason, et al.
Published: (2025)
BAGELS: Benchmarking the Automated Generation and Extraction of Limitations from Scholarly Text
by: Azher, Ibrahim Al, et al.
Published: (2025)
by: Azher, Ibrahim Al, et al.
Published: (2025)
L-PRISMA: An Extension of PRISMA in the Era of Generative Artificial Intelligence (GenAI)
by: Shailendra, Samar, et al.
Published: (2026)
by: Shailendra, Samar, et al.
Published: (2026)
Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
by: Rondina, Marco, et al.
Published: (2025)
by: Rondina, Marco, et al.
Published: (2025)
Public interest in science or bots? Selective amplification of scientific articles on Twitter
by: Rahman, Ashiqur, et al.
Published: (2024)
by: Rahman, Ashiqur, et al.
Published: (2024)
Persistence Paradox in Dynamic Science
by: Bao, Honglin, et al.
Published: (2025)
by: Bao, Honglin, et al.
Published: (2025)
Automating Violence Detection and Categorization from Ancient Texts
by: Abdelhalim, Alhassan, et al.
Published: (2025)
by: Abdelhalim, Alhassan, et al.
Published: (2025)
Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
by: Liu, Fan, et al.
Published: (2025)
by: Liu, Fan, et al.
Published: (2025)
On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts
by: Shahi, Gautam Kishore, et al.
Published: (2025)
by: Shahi, Gautam Kishore, et al.
Published: (2025)
Similar Items
-
Leveraging Expert Input for Robust and Explainable AI-Assisted Lung Cancer Detection in Chest X-rays
by: Rafferty, Amy, et al.
Published: (2024) -
Simbanex: Similarity-based Exploration of IEEE VIS Publications
by: Witschard, Daniel, et al.
Published: (2024) -
FedAIoT: A Federated Learning Benchmark for Artificial Intelligence of Things
by: Alam, Samiul, et al.
Published: (2023) -
Sanidha: A Studio Quality Multi-Modal Dataset for Carnatic Music
by: Krishnan, Venkatakrishnan Vaidyanathapuram, et al.
Published: (2025) -
ARED: Argentina Real Estate Dataset
by: Belenky, Iván
Published: (2024)