AI-generated data contamination erodes pathological variability and diagnostic reliability
Fuente:
arXiv
Saved in:
| Main Authors: | He, Hongyu, Xiang, Shaowen, Zhang, Ye, Zhu, Yingtao, Zhang, Jin, Deng, Hao, Alsentzer, Emily, Liu, Yun, Chen, Qingyu, Yu, Kun-Hsing, Marshall, Andrew, Chen, Tingting, Anumasa, Srinivas, Ebner, Daniel, Ho, Dean, Ngiam, Kee Yuan, Cheng, Ching-Yu, Liu, Dianbo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via inference-time Multi-Views Brainstorming
by: Lagzian, Arash, et al.
Published: (2025)
by: Lagzian, Arash, et al.
Published: (2025)
Laplacian Score Sharpening for Mitigating Hallucination in Diffusion Models
by: C, Barath Chandran., et al.
Published: (2025)
by: C, Barath Chandran., et al.
Published: (2025)
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
by: Anumasa, Srinivas, et al.
Published: (2025)
by: Anumasa, Srinivas, et al.
Published: (2025)
Navigating heterogeneous protein landscapes through geometry-aware smoothing
by: Anumasa, Srinivas, et al.
Published: (2026)
by: Anumasa, Srinivas, et al.
Published: (2026)
HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
by: Chen, Tingting, et al.
Published: (2025)
by: Chen, Tingting, et al.
Published: (2025)
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
by: Wang, Yixuan, et al.
Published: (2026)
by: Wang, Yixuan, et al.
Published: (2026)
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
by: Zou, Qiran, et al.
Published: (2025)
by: Zou, Qiran, et al.
Published: (2025)
Safety challenges of AI in medicine in the era of large language models
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
by: Yuan, Qianhao, et al.
Published: (2025)
by: Yuan, Qianhao, et al.
Published: (2025)
CodeUnlearn: Amortized Zero-Shot Machine Unlearning in Language Models Using Discrete Concept
by: Wu, YuXuan, et al.
Published: (2024)
by: Wu, YuXuan, et al.
Published: (2024)
Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
by: Weishaupt, Luca L., et al.
Published: (2025)
by: Weishaupt, Luca L., et al.
Published: (2025)
River-LLM: Large Language Model Seamless Exit Based on KV Share
by: Shen, Yingtao, et al.
Published: (2026)
by: Shen, Yingtao, et al.
Published: (2026)
DiffCJK: Conditional Diffusion Model for High-Quality and Wide-coverage CJK Character Generation
by: Tian, Yingtao
Published: (2024)
by: Tian, Yingtao
Published: (2024)
Balance of Number of Embedding and their Dimensions in Vector Quantization
by: Chen, Hang, et al.
Published: (2024)
by: Chen, Hang, et al.
Published: (2024)
Base of RoPE Bounds Context Length
by: Men, Xin, et al.
Published: (2024)
by: Men, Xin, et al.
Published: (2024)
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
by: Chen, Tingting, et al.
Published: (2025)
by: Chen, Tingting, et al.
Published: (2025)
SURE: SUrvey REcipes for building reliable and robust deep networks
by: Li, Yuting, et al.
Published: (2024)
by: Li, Yuting, et al.
Published: (2024)
LoRA-GA: Low-Rank Adaptation with Gradient Approximation
by: Wang, Shaowen, et al.
Published: (2024)
by: Wang, Shaowen, et al.
Published: (2024)
Intimate Partner Violence and Injury Prediction From Radiology Reports
by: Chen, Irene Y., et al.
Published: (2020)
by: Chen, Irene Y., et al.
Published: (2020)
Smartphone-Based 3D Indoor Localization and Navigation (Volume 1)
by: Ebner, Frank
Published: (2022)
by: Ebner, Frank
Published: (2022)
Emotional Supporters often Use Multiple Strategies in a Single Turn
by: Bai, Xin, et al.
Published: (2025)
by: Bai, Xin, et al.
Published: (2025)
Generative AIBIM: An automatic and intelligent structural design pipeline integrating BIM and generative AI
by: He, Zhili, et al.
Published: (2023)
by: He, Zhili, et al.
Published: (2023)
Large language models eroding science understanding: an experimental study
by: Collins, Harry, et al.
Published: (2026)
by: Collins, Harry, et al.
Published: (2026)
The impact of tissue detection on diagnostic artificial intelligence algorithms in digital pathology
by: Boman, Sol Erika, et al.
Published: (2025)
by: Boman, Sol Erika, et al.
Published: (2025)
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
by: Zou, Qiran, et al.
Published: (2026)
by: Zou, Qiran, et al.
Published: (2026)
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
by: Men, Xin, et al.
Published: (2024)
by: Men, Xin, et al.
Published: (2024)
Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models
by: Tian, Bozhong, et al.
Published: (2024)
by: Tian, Bozhong, et al.
Published: (2024)
MedINST: Meta Dataset of Biomedical Instructions
by: Han, Wenhan, et al.
Published: (2024)
by: Han, Wenhan, et al.
Published: (2024)
Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
CXR-LanIC: Language-Grounded Interpretable Classifier for Chest X-Ray Diagnosis
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Spectral Complex Autoencoder Pruning: A Fidelity-Guided Criterion for Extreme Structured Channel Compression
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Mamba meets crack segmentation
by: He, Zhili, et al.
Published: (2024)
by: He, Zhili, et al.
Published: (2024)
DDIM sampling for Generative AIBIM, a faster intelligent structural design framework
by: He, Zhili, et al.
Published: (2024)
by: He, Zhili, et al.
Published: (2024)
Scaling Ultrasound Volumetric Reconstruction via Mobile Augmented Reality
by: Ng, Kian Wei, et al.
Published: (2026)
by: Ng, Kian Wei, et al.
Published: (2026)
Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers
by: Xin, Rihui, et al.
Published: (2025)
by: Xin, Rihui, et al.
Published: (2025)
LLMSR@XLLM25: An Empirical Study of LLM for Structural Reasoning
by: Li, Xinye, et al.
Published: (2025)
by: Li, Xinye, et al.
Published: (2025)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Similar Items
-
Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via inference-time Multi-Views Brainstorming
by: Lagzian, Arash, et al.
Published: (2025) -
Laplacian Score Sharpening for Mitigating Hallucination in Diffusion Models
by: C, Barath Chandran., et al.
Published: (2025) -
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
by: Tang, Yiming, et al.
Published: (2025) -
Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
by: Anumasa, Srinivas, et al.
Published: (2025) -
Navigating heterogeneous protein landscapes through geometry-aware smoothing
by: Anumasa, Srinivas, et al.
Published: (2026)