DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
Fuente:
arXiv
Saved in:
| Main Author: | Oleson, Jon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning
by: Jin, Congyun, et al.
Published: (2024)
by: Jin, Congyun, et al.
Published: (2024)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
by: Devunuri, Saipraneeth, et al.
Published: (2024)
by: Devunuri, Saipraneeth, et al.
Published: (2024)
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
by: Hu, Zhengmian, et al.
Published: (2024)
by: Hu, Zhengmian, et al.
Published: (2024)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024)
by: Hardy, Michael
Published: (2024)
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages
by: Schöffel, Matthias, et al.
Published: (2026)
by: Schöffel, Matthias, et al.
Published: (2026)
Language-Dependent Political Bias in AI: A Study of ChatGPT and Gemini
by: Yuksel, Dogus, et al.
Published: (2025)
by: Yuksel, Dogus, et al.
Published: (2025)
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Improving LLM Leaderboards with Psychometrical Methodology
by: Federiakin, Denis
Published: (2025)
by: Federiakin, Denis
Published: (2025)
AI-Assisted Decision-Making for Clinical Assessment of Auto-Segmented Contour Quality
by: Wang, Biling, et al.
Published: (2025)
by: Wang, Biling, et al.
Published: (2025)
United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
by: von der Heyde, Leah, et al.
Published: (2024)
by: von der Heyde, Leah, et al.
Published: (2024)
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Metacognitive Myopia in Large Language Models
by: Scholten, Florian, et al.
Published: (2024)
by: Scholten, Florian, et al.
Published: (2024)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
by: Hang, Haotian, et al.
Published: (2025)
by: Hang, Haotian, et al.
Published: (2025)
A Rational Analysis of the Speech-to-Song Illusion
by: Marjieh, Raja, et al.
Published: (2024)
by: Marjieh, Raja, et al.
Published: (2024)
ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational Theory
by: Yuan, Yunhao, et al.
Published: (2025)
by: Yuan, Yunhao, et al.
Published: (2025)
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
by: Yuan, Hongyi, et al.
Published: (2024)
by: Yuan, Hongyi, et al.
Published: (2024)
Augmented Risk Prediction for the Onset of Alzheimer's Disease from Electronic Health Records with Large Language Models
by: Wang, Jiankun, et al.
Published: (2024)
by: Wang, Jiankun, et al.
Published: (2024)
Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
by: Kokkodis, Marios, et al.
Published: (2025)
by: Kokkodis, Marios, et al.
Published: (2025)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
by: Tschisgale, Paul, et al.
Published: (2026)
by: Tschisgale, Paul, et al.
Published: (2026)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
by: Jantre, Sanket, et al.
Published: (2025)
by: Jantre, Sanket, et al.
Published: (2025)
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
by: Dobbins, Nic, et al.
Published: (2025)
by: Dobbins, Nic, et al.
Published: (2025)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Crowdsourced Adaptive Surveys
by: Velez, Yamil
Published: (2024)
by: Velez, Yamil
Published: (2024)
Limits of Large Language Models in Debating Humans
by: Flamino, James, et al.
Published: (2024)
by: Flamino, James, et al.
Published: (2024)
AI-Assisted Conversational Interviewing: Effects on Data Quality and Respondent Experience
by: Barari, Soubhik, et al.
Published: (2025)
by: Barari, Soubhik, et al.
Published: (2025)
Generative AI Spotlights the Human Core of Data Science: Implications for Education
by: Taback, Nathan
Published: (2026)
by: Taback, Nathan
Published: (2026)
The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control
by: Lommel, Arle, et al.
Published: (2024)
by: Lommel, Arle, et al.
Published: (2024)
ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms
by: Vail, Alexandria K., et al.
Published: (2026)
by: Vail, Alexandria K., et al.
Published: (2026)
Depression Detection Using Digital Traces on Social Media: A Knowledge-aware Deep Learning Approach
by: Zhang, Wenli, et al.
Published: (2023)
by: Zhang, Wenli, et al.
Published: (2023)
Removing Spurious Correlation from Neural Network Interpretations
by: Fotouhi, Milad, et al.
Published: (2024)
by: Fotouhi, Milad, et al.
Published: (2024)
Bridging the Data Gap in AI Reliability Research and Establishing DR-AIR, a Comprehensive Data Repository for AI Reliability
by: Zheng, Simin, et al.
Published: (2025)
by: Zheng, Simin, et al.
Published: (2025)
"Crash Test Dummies" for AI-Enabled Clinical Assessment: Validating Virtual Patient Scenarios with Virtual Learners
by: Gin, Brian, et al.
Published: (2026)
by: Gin, Brian, et al.
Published: (2026)
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement
by: Sielinski, Ronald
Published: (2026)
by: Sielinski, Ronald
Published: (2026)
Similar Items
-
RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning
by: Jin, Congyun, et al.
Published: (2024) -
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
by: Devunuri, Saipraneeth, et al.
Published: (2024) -
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
by: Hu, Zhengmian, et al.
Published: (2024) -
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024) -
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)