Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
Fuente:
arXiv
Saved in:
| Main Authors: | Flores, Gerardo A., Smith, Alyssa H., Fukuyama, Julia A., Wilson, Ashia C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools
by: Flores, Gerardo, et al.
Published: (2025)
by: Flores, Gerardo, et al.
Published: (2025)
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
The Approximate Fisher Influence Function: Faster Estimation of Data Influence in Statistical Models
by: Lev, Omri, et al.
Published: (2024)
by: Lev, Omri, et al.
Published: (2024)
Low-Cost Labels, Reliable Choices: Rollout-Calibrated Hyper-Heuristics for Job Shop Scheduling
by: Wei, Junhao, et al.
Published: (2026)
by: Wei, Junhao, et al.
Published: (2026)
Calibrated Preference Learning: The Case of Label Ranking
by: Thies, Santo M. A. R., et al.
Published: (2026)
by: Thies, Santo M. A. R., et al.
Published: (2026)
The Cost of Relaxation: Evaluating the Error in Convex Neural Network Verification
by: Papamichail, Merkouris, et al.
Published: (2026)
by: Papamichail, Merkouris, et al.
Published: (2026)
AdapTable: Test-Time Adaptation for Tabular Data via Shift-Aware Uncertainty Calibrator and Label Distribution Handler
by: Kim, Changhun, et al.
Published: (2024)
by: Kim, Changhun, et al.
Published: (2024)
Calibration Error Estimation Using Fuzzy Binning
by: Bihani, Geetanjali, et al.
Published: (2023)
by: Bihani, Geetanjali, et al.
Published: (2023)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Pandora's Regret: A Proper Scoring Rule for Evaluating Sequential Search
by: Flores, Gerardo A., et al.
Published: (2026)
by: Flores, Gerardo A., et al.
Published: (2026)
Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
by: Xiao, Jiancong, et al.
Published: (2025)
by: Xiao, Jiancong, et al.
Published: (2025)
Dirichlet-Based Prediction Calibration for Learning with Noisy Labels
by: Zong, Chen-Chen, et al.
Published: (2024)
by: Zong, Chen-Chen, et al.
Published: (2024)
Robust Calibration For Improved Weather Prediction Under Distributional Shift
by: Gilda, Sankalp, et al.
Published: (2024)
by: Gilda, Sankalp, et al.
Published: (2024)
Dynamics-Aligned Shared Hypernetworks for Contextual RL under Discontinuous Shifts
by: Benad, Jan, et al.
Published: (2026)
by: Benad, Jan, et al.
Published: (2026)
Improving Label Error Detection and Elimination with Uncertainty Quantification
by: Jakubik, Johannes, et al.
Published: (2024)
by: Jakubik, Johannes, et al.
Published: (2024)
Addressing Label Shift in Distributed Learning via Entropy Regularization
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Dynamical Label Augmentation and Calibration for Noisy Electronic Health Records
by: Li, Yuhao, et al.
Published: (2025)
by: Li, Yuhao, et al.
Published: (2025)
Calibrated Language Models and How to Find Them with Label Smoothing
by: Huang, Jerry, et al.
Published: (2025)
by: Huang, Jerry, et al.
Published: (2025)
CLUE: Neural Networks Calibration via Learning Uncertainty-Error alignment
by: Mendes, Pedro, et al.
Published: (2025)
by: Mendes, Pedro, et al.
Published: (2025)
Calibration of Time-Series Forecasting: Detecting and Adapting Context-Driven Distribution Shift
by: Chen, Mouxiang, et al.
Published: (2023)
by: Chen, Mouxiang, et al.
Published: (2023)
Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
by: Kusaka, Shigeki, et al.
Published: (2025)
by: Kusaka, Shigeki, et al.
Published: (2025)
Mind the GAP: Improving Robustness to Subpopulation Shifts with Group-Aware Priors
by: Rudner, Tim G. J., et al.
Published: (2024)
by: Rudner, Tim G. J., et al.
Published: (2024)
Class-based Subset Selection for Transfer Learning under Extreme Label Shift
by: Goyal, Akul, et al.
Published: (2024)
by: Goyal, Akul, et al.
Published: (2024)
Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
by: De Simone, Zoe, et al.
Published: (2023)
by: De Simone, Zoe, et al.
Published: (2023)
Correcting Noisy Multilabel Predictions: Modeling Label Noise through Latent Space Shifts
by: Huang, Weipeng, et al.
Published: (2025)
by: Huang, Weipeng, et al.
Published: (2025)
Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior Adjustment
by: Lee, Seunghan
Published: (2026)
by: Lee, Seunghan
Published: (2026)
SEF: A Method for Computing Prediction Intervals by Shifting the Error Function in Neural Networks
by: Aretos, E. V., et al.
Published: (2024)
by: Aretos, E. V., et al.
Published: (2024)
FAQ: Mitigating Quantization Error via Regenerating Calibration Data with Family-Aware Quantization
by: Xiao, Haiyang, et al.
Published: (2026)
by: Xiao, Haiyang, et al.
Published: (2026)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
by: Dabas, Mahavir, et al.
Published: (2025)
by: Dabas, Mahavir, et al.
Published: (2025)
UniAlign: A Model-Agnostic Framework for Robust Network Traffic Classification under Distribution Shifts
by: Wang, Tongze, et al.
Published: (2026)
by: Wang, Tongze, et al.
Published: (2026)
Generalized Priority-Aware Shapley Value
by: Lee, Kiljae, et al.
Published: (2026)
by: Lee, Kiljae, et al.
Published: (2026)
Quantifying Calibration Error in Neural Networks Through Evidence-Based Theory
by: Ouattara, Koffi Ismael, et al.
Published: (2024)
by: Ouattara, Koffi Ismael, et al.
Published: (2024)
Exact Stiefel Optimization for Probabilistic PLS: Closed-Form Updates, Error Bounds, and Calibrated Uncertainty
by: Hu, Haoran, et al.
Published: (2026)
by: Hu, Haoran, et al.
Published: (2026)
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
by: Nguyen, Thanh Thi, et al.
Published: (2025)
by: Nguyen, Thanh Thi, et al.
Published: (2025)
HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees
by: Zeng, Hao, et al.
Published: (2026)
by: Zeng, Hao, et al.
Published: (2026)
Rethinking Soft Actor-Critic in High-Dimensional Action Spaces: The Cost of Ignoring Distribution Shift
by: Chen, Yanjun, et al.
Published: (2024)
by: Chen, Yanjun, et al.
Published: (2024)
Even-if Explanations: Formal Foundations, Priorities and Complexity
by: Alfano, Gianvincenzo, et al.
Published: (2024)
by: Alfano, Gianvincenzo, et al.
Published: (2024)
On the Usefulness of the Fit-on-the-Test View on Evaluating Calibration of Classifiers
by: Kängsepp, Markus, et al.
Published: (2022)
by: Kängsepp, Markus, et al.
Published: (2022)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
Similar Items
-
A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools
by: Flores, Gerardo, et al.
Published: (2025) -
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025) -
The Approximate Fisher Influence Function: Faster Estimation of Data Influence in Statistical Models
by: Lev, Omri, et al.
Published: (2024) -
Low-Cost Labels, Reliable Choices: Rollout-Calibrated Hyper-Heuristics for Job Shop Scheduling
by: Wei, Junhao, et al.
Published: (2026) -
Calibrated Preference Learning: The Case of Label Ranking
by: Thies, Santo M. A. R., et al.
Published: (2026)