A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools
Fuente:
arXiv
Saved in:
| Main Authors: | Flores, Gerardo, Schiff, Abigail, Smith, Alyssa H., Fukuyama, Julia A, Wilson, Ashia C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
by: Flores, Gerardo A., et al.
Published: (2025)
by: Flores, Gerardo A., et al.
Published: (2025)
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
SHAP-Based Supervised Clustering for Sample Classification and the Generalized Waterfall Plot
by: Lin, Justin, et al.
Published: (2025)
by: Lin, Justin, et al.
Published: (2025)
Consequentialist Objectives and Catastrophe
by: Marklund, Henrik, et al.
Published: (2026)
by: Marklund, Henrik, et al.
Published: (2026)
Mind the GAP: Improving Robustness to Subpopulation Shifts with Group-Aware Priors
by: Rudner, Tim G. J., et al.
Published: (2024)
by: Rudner, Tim G. J., et al.
Published: (2024)
The Approximate Fisher Influence Function: Faster Estimation of Data Influence in Statistical Models
by: Lev, Omri, et al.
Published: (2024)
by: Lev, Omri, et al.
Published: (2024)
Exploring Multi-Modal Data with Tool-Augmented LLM Agents for Precise Causal Discovery
by: Shen, ChengAo, et al.
Published: (2024)
by: Shen, ChengAo, et al.
Published: (2024)
Causal and Local Correlations Based Network for Multivariate Time Series Classification
by: Du, Mingsen, et al.
Published: (2024)
by: Du, Mingsen, et al.
Published: (2024)
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
by: Xu, Zhiyu, et al.
Published: (2025)
by: Xu, Zhiyu, et al.
Published: (2025)
Robust Classification of High-Dimensional Data using Data-Adaptive Energy Distance
by: Choudhury, Jyotishka Ray, et al.
Published: (2023)
by: Choudhury, Jyotishka Ray, et al.
Published: (2023)
A Critical Perspective on Finite Sample Conformal Prediction Theory in Medical Applications
by: Kladny, Klaus-Rudolf, et al.
Published: (2025)
by: Kladny, Klaus-Rudolf, et al.
Published: (2025)
Teleporter Theory: A General and Simple Approach for Modeling Cross-World Counterfactual Causality
by: Li, Jiangmeng, et al.
Published: (2024)
by: Li, Jiangmeng, et al.
Published: (2024)
Evaluating the Effectiveness of Index-Based Treatment Allocation
by: Boehmer, Niclas, et al.
Published: (2024)
by: Boehmer, Niclas, et al.
Published: (2024)
A Causal Framework for Evaluating ICU Discharge Strategies
by: Simha, Sagar Nagaraj, et al.
Published: (2026)
by: Simha, Sagar Nagaraj, et al.
Published: (2026)
Weak Supervision Performance Evaluation via Partial Identification
by: Polo, Felipe Maia, et al.
Published: (2023)
by: Polo, Felipe Maia, et al.
Published: (2023)
Characterising harmful data sources when constructing multi-fidelity surrogate models
by: Andrés-Thió, Nicolau, et al.
Published: (2024)
by: Andrés-Thió, Nicolau, et al.
Published: (2024)
Off-Policy Evaluation and Learning for Survival Outcomes under Censoring
by: Kubota, Kohsuke, et al.
Published: (2026)
by: Kubota, Kohsuke, et al.
Published: (2026)
CausalCompass: Evaluating the Robustness of Time-Series Causal Discovery in Misspecified Scenarios
by: Yi, Huiyang, et al.
Published: (2026)
by: Yi, Huiyang, et al.
Published: (2026)
ProCause: Generating Counterfactual Outcomes to Evaluate Prescriptive Process Monitoring Methods
by: De Moor, Jakob, et al.
Published: (2025)
by: De Moor, Jakob, et al.
Published: (2025)
Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation
by: Martinon, Grégoire, et al.
Published: (2026)
by: Martinon, Grégoire, et al.
Published: (2026)
Evaluation of Stress Detection as Time Series Events -- A Novel Window-Based F1-Metric
by: Skat-Rørdam, Harald Vilhelm, et al.
Published: (2025)
by: Skat-Rørdam, Harald Vilhelm, et al.
Published: (2025)
Conformal Tradeoffs: Operational Profiles Beyond Coverage
by: Zwart, Petrus H.
Published: (2026)
by: Zwart, Petrus H.
Published: (2026)
A Two-Stage Feature Selection Approach for Robust Evaluation of Treatment Effects in High-Dimensional Observational Data
by: Islam, Md Saiful, et al.
Published: (2021)
by: Islam, Md Saiful, et al.
Published: (2021)
A weighted U statistic for association analysis considering genetic heterogeneity
by: Wei, Changshuai, et al.
Published: (2015)
by: Wei, Changshuai, et al.
Published: (2015)
A Data-Driven Two-Phase Multi-Split Causal Ensemble Model for Time Series
by: Ma, Zhipeng, et al.
Published: (2024)
by: Ma, Zhipeng, et al.
Published: (2024)
Bayesian Hierarchical Invariant Prediction
by: Madaleno, Francisco, et al.
Published: (2025)
by: Madaleno, Francisco, et al.
Published: (2025)
A Generalized Genetic Random Field Method for the Genetic Association Analysis of Sequencing Data
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Toward a Theory of Causation for Interpreting Neural Code Models
by: Palacio, David N., et al.
Published: (2023)
by: Palacio, David N., et al.
Published: (2023)
Seeing Through Risk: A Symbolic Approximation of Prospect Theory
by: Yousaf, Ali Arslan, et al.
Published: (2025)
by: Yousaf, Ali Arslan, et al.
Published: (2025)
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
by: Dong, Zihan, et al.
Published: (2026)
by: Dong, Zihan, et al.
Published: (2026)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
by: Keith, Katherine A., et al.
Published: (2023)
by: Keith, Katherine A., et al.
Published: (2023)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
Econometric vs. Causal Structure-Learning for Time-Series Policy Decisions: Evidence from the UK COVID-19 Policies
by: Petrungaro, Bruno, et al.
Published: (2026)
by: Petrungaro, Bruno, et al.
Published: (2026)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Worst-case low-rank approximations
by: Fries, Anya, et al.
Published: (2026)
by: Fries, Anya, et al.
Published: (2026)
rSDNet: Unified Robust Neural Learning against Label Noise and Adversarial Attacks
by: Jana, Suryasis, et al.
Published: (2026)
by: Jana, Suryasis, et al.
Published: (2026)
Enhancing Maritime Trajectory Forecasting via H3 Index and Causal Language Modelling (CLM)
by: Drapier, Nicolas, et al.
Published: (2024)
by: Drapier, Nicolas, et al.
Published: (2024)
A multi-locus predictiveness curve and its summary assessment for genetic risk prediction
by: Wei, Changshuai, et al.
Published: (2025)
by: Wei, Changshuai, et al.
Published: (2025)
Effective Bayesian Causal Inference via Structural Marginalisation and Autoregressive Orders
by: Toth, Christian, et al.
Published: (2024)
by: Toth, Christian, et al.
Published: (2024)
Similar Items
-
Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
by: Flores, Gerardo A., et al.
Published: (2025) -
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025) -
SHAP-Based Supervised Clustering for Sample Classification and the Generalized Waterfall Plot
by: Lin, Justin, et al.
Published: (2025) -
Consequentialist Objectives and Catastrophe
by: Marklund, Henrik, et al.
Published: (2026) -
Mind the GAP: Improving Robustness to Subpopulation Shifts with Group-Aware Priors
by: Rudner, Tim G. J., et al.
Published: (2024)