Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
Fuente:
arXiv
Saved in:
| Main Authors: | Maslenkova, Svetlana, Christophe, Clement, Pimentel, Marco AF, Raha, Tathagata, Salman, Muhammad Umar, Mahrooqi, Ahmed Al, Gupta, Avani, Khan, Shadab, Rajan, Ronnie, Kanithi, Praveenkumar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs
by: Christophe, Clément, et al.
Published: (2024)
by: Christophe, Clément, et al.
Published: (2024)
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
by: Kanithi, Praveenkumar, et al.
Published: (2024)
by: Kanithi, Praveenkumar, et al.
Published: (2024)
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation
by: Raha, Tathagata, et al.
Published: (2026)
by: Raha, Tathagata, et al.
Published: (2026)
Med42-v2: A Suite of Clinical LLMs
by: Christophe, Clément, et al.
Published: (2024)
by: Christophe, Clément, et al.
Published: (2024)
Named Clinical Entity Recognition Benchmark
by: Abdul, Wadood M, et al.
Published: (2024)
by: Abdul, Wadood M, et al.
Published: (2024)
Bridging Language Barriers in Healthcare: A Study on Arabic LLMs
by: Saadi, Nada, et al.
Published: (2025)
by: Saadi, Nada, et al.
Published: (2025)
Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare
by: Christophe, Clément, et al.
Published: (2026)
by: Christophe, Clément, et al.
Published: (2026)
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
by: Pimentel, Marco AF, et al.
Published: (2024)
by: Pimentel, Marco AF, et al.
Published: (2024)
Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches
by: Christophe, Clément, et al.
Published: (2024)
by: Christophe, Clément, et al.
Published: (2024)
Do Instruction-Tuned Models Always Perform Better Than Base Models? Evidence from Math and Domain-Shifted Benchmarks
by: Munjal, Prateek, et al.
Published: (2026)
by: Munjal, Prateek, et al.
Published: (2026)
iREL at SemEval-2024 Task 9: Improving Conventional Prompting Methods for Brain Teasers
by: Gupta, Harshit, et al.
Published: (2024)
by: Gupta, Harshit, et al.
Published: (2024)
Gene42: Long-Range Genomic Foundation Model With Dense Attention
by: Vishniakov, Kirill, et al.
Published: (2025)
by: Vishniakov, Kirill, et al.
Published: (2025)
Building Trust: Foundations of Security, Safety and Transparency in AI
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
A survey on Concept-based Approaches For Model Improvement
by: Gupta, Avani, et al.
Published: (2024)
by: Gupta, Avani, et al.
Published: (2024)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
by: Malik, Hashmat Shadab, et al.
Published: (2025)
by: Malik, Hashmat Shadab, et al.
Published: (2025)
Cross-Country Associations of the Happiness Index
by: Khosla, Avani
Published: (2025)
by: Khosla, Avani
Published: (2025)
Prediction on mechanical properties of engineered cementitious composites: An experimental and machine learning approach
by: N. Shanmugasundaram, et al.
Published: (2024)
by: N. Shanmugasundaram, et al.
Published: (2024)
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
by: Malik, Hashmat Shadab, et al.
Published: (2025)
by: Malik, Hashmat Shadab, et al.
Published: (2025)
Bridging the Data Gap: Creating a Hindi Text Summarization Dataset from the English XSUM
by: Katwe, Praveenkumar, et al.
Published: (2026)
by: Katwe, Praveenkumar, et al.
Published: (2026)
Quran-MD: A Fine-Grained Multilingual Multimodal Dataset of the Quran
by: Salman, Muhammad Umar, et al.
Published: (2026)
by: Salman, Muhammad Umar, et al.
Published: (2026)
Towards Evaluating the Robustness of Visual State Space Models
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Limitations of Public Chest Radiography Datasets for Artificial Intelligence: Label Quality, Domain Shift, Bias and Evaluation Challenges
by: Rafferty, Amy, et al.
Published: (2025)
by: Rafferty, Amy, et al.
Published: (2025)
Três artes no Teatro de Reprise em um psicodrama público
by: Rosane Avani Rodrigues
Published: (2017)
by: Rosane Avani Rodrigues
Published: (2017)
Enhancement of Solid Particle Erosion‐Resistance in Carbon‐Fiber Epoxy Composites Using Electrophoretically Deposited Carboxyl Functionalized Graphene on Carbon Fiber
by: Praveenkumar Jatothu, et al.
Published: (2026)
by: Praveenkumar Jatothu, et al.
Published: (2026)
Electrophoretic Deposition of Nanoparticles on Carbon Fiber: A Comprehensive Review of Enhancement of Properties in Epoxy Polymer Composites
by: Praveenkumar Jatothu, et al.
Published: (2025)
by: Praveenkumar Jatothu, et al.
Published: (2025)
Flowing Datasets with Wasserstein over Wasserstein Gradient Flows
by: Bonet, Clément, et al.
Published: (2025)
by: Bonet, Clément, et al.
Published: (2025)
CLEANANERCorp: Identifying and Correcting Incorrect Labels in the ANERcorp Dataset
by: Al-Duwais, Mashael, et al.
Published: (2024)
by: Al-Duwais, Mashael, et al.
Published: (2024)
Trust and Transparency in an Age of Surveillance
Published: (2021)
Published: (2021)
A wide voltage range non‐isolated continuous input buck‐boost converter for optimal green energy harvesting in LED lighting systems
by: Ashok Kumar Kanithi, et al.
Published: (2024)
by: Ashok Kumar Kanithi, et al.
Published: (2024)
Performance analysis of a bridgeless power factor correction (PFC) buck–boost LED driver with ripple diversion scheme for extended lifespan
by: Kanithi Ashok Kumar, et al.
Published: (2024)
by: Kanithi Ashok Kumar, et al.
Published: (2024)
Overcoming Ambient Drift and Negative-Bias Temperature Instability in Foundry Carbon Nanotube Transistors
by: Yu, Andrew, et al.
Published: (2024)
by: Yu, Andrew, et al.
Published: (2024)
On Evaluating Adversarial Robustness of Volumetric Medical Segmentation Models
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Isovector Axial Charge and Form Factors of Nucleons from Lattice QCD
by: Gupta, Rajan
Published: (2024)
by: Gupta, Rajan
Published: (2024)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
Fundamentals of Next-generation Network Planning
by: Khan, M. Umar
Published: (2025)
by: Khan, M. Umar
Published: (2025)
Transforming Next-generation Network Planning assisted by Data Acquisition of Top Three Spanish MNOs
by: Khan, M. Umar
Published: (2025)
by: Khan, M. Umar
Published: (2025)
A Confidential Computing Transparency Framework for a Comprehensive Trust Chain
by: Kocaoğullar, Ceren, et al.
Published: (2024)
by: Kocaoğullar, Ceren, et al.
Published: (2024)
Unambiguous discrimination of sequences of quantum states
by: Gupta, Tathagata, et al.
Published: (2024)
by: Gupta, Tathagata, et al.
Published: (2024)
The Future of ChatGPT in Medicinal Chemistry: Harnessing AI for Accelerated Drug Discovery
by: Tathagata Pradhan, et al.
Published: (2024)
by: Tathagata Pradhan, et al.
Published: (2024)
Similar Items
-
Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs
by: Christophe, Clément, et al.
Published: (2024) -
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
by: Kanithi, Praveenkumar, et al.
Published: (2024) -
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation
by: Raha, Tathagata, et al.
Published: (2026) -
Med42-v2: A Suite of Clinical LLMs
by: Christophe, Clément, et al.
Published: (2024) -
Named Clinical Entity Recognition Benchmark
by: Abdul, Wadood M, et al.
Published: (2024)