Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings
Fuente:
arXiv
Saved in:
| Main Authors: | Pouget, Angéline, Yaghini, Mohammad, Rabanser, Stephan, Papernot, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Does It Take to Build a Performant Selective Classifier?
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
by: Rabanser, Stephan
Published: (2025)
by: Rabanser, Stephan
Published: (2025)
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026)
by: Rabanser, Stephan, et al.
Published: (2026)
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
by: Patwardhan, Tejal, et al.
Published: (2025)
by: Patwardhan, Tejal, et al.
Published: (2025)
Reinforced Sequential Decision-Making for Sepsis Treatment: The POSNEGDM Framework with Mortality Classifier and Transformer
by: Tamboli, Dipesh, et al.
Published: (2024)
by: Tamboli, Dipesh, et al.
Published: (2024)
FairJob: A Real-World Dataset for Fairness in Online Systems
by: Vladimirova, Mariia, et al.
Published: (2024)
by: Vladimirova, Mariia, et al.
Published: (2024)
Regulation Games for Trustworthy Machine Learning
by: Yaghini, Mohammad, et al.
Published: (2024)
by: Yaghini, Mohammad, et al.
Published: (2024)
When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital
by: Kothari, Avni, et al.
Published: (2025)
by: Kothari, Avni, et al.
Published: (2025)
Back to the Drawing Board for Fair Representation Learning
by: Pouget, Angéline, et al.
Published: (2024)
by: Pouget, Angéline, et al.
Published: (2024)
A Regulatory Governance Framework for AI-Driven Financial Fraud Detection in U.S. Banking: Integrating OCC, SR 11-7, CFPB, and FinCEN Compliance Requirements for Model Development, Validation, and Monitoring Lifecycles
by: Uddin, Mohammad Nasir
Published: (2026)
by: Uddin, Mohammad Nasir
Published: (2026)
Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes
by: Toups, Connor, et al.
Published: (2023)
by: Toups, Connor, et al.
Published: (2023)
Beyond Algorithmic Fairness: A Guide to Develop and Deploy Ethical AI-Enabled Decision-Support Tools
by: Gonzalez, Rosemarie Santa, et al.
Published: (2024)
by: Gonzalez, Rosemarie Santa, et al.
Published: (2024)
Developing and Deploying Industry Standards for Artificial Intelligence in Education (AIED): Challenges, Strategies, and Future Directions
by: Tong, Richard, et al.
Published: (2024)
by: Tong, Richard, et al.
Published: (2024)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
by: Huang, Saffron, et al.
Published: (2025)
by: Huang, Saffron, et al.
Published: (2025)
OPTIC-ER: A Reinforcement Learning Framework for Real-Time Emergency Response and Equitable Resource Allocation in Underserved African Communities
by: Tonwe, Mary
Published: (2025)
by: Tonwe, Mary
Published: (2025)
AI-powered Digital Framework for Personalized Economical Quality Learning at Scale
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
by: Marius, Dumitran Adrian, et al.
Published: (2025)
by: Marius, Dumitran Adrian, et al.
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
An Epistemic and Aleatoric Decomposition of Arbitrariness to Constrain the Set of Good Models
by: Khan, Falaah Arif, et al.
Published: (2023)
by: Khan, Falaah Arif, et al.
Published: (2023)
Optimizing HIV Patient Engagement with Reinforcement Learning in Resource-Limited Settings
by: Periáñez, África, et al.
Published: (2024)
by: Periáñez, África, et al.
Published: (2024)
From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution
by: Koch, Bernard J., et al.
Published: (2024)
by: Koch, Bernard J., et al.
Published: (2024)
What Is Fairness? On the Role of Protected Attributes and Fictitious Worlds
by: Bothmann, Ludwig, et al.
Published: (2022)
by: Bothmann, Ludwig, et al.
Published: (2022)
Rethinking Suicidal Ideation Detection: A Trustworthy Annotation Framework and Cross-Lingual Model Evaluation
by: Dzafic, Amina, et al.
Published: (2025)
by: Dzafic, Amina, et al.
Published: (2025)
FairHealth: An Open-Source Python Library for Trustworthy Healthcare AI in Low-Resource Settings
by: Yesmin, Farjana
Published: (2026)
by: Yesmin, Farjana
Published: (2026)
Clio: Privacy-Preserving Insights into Real-World AI Use
by: Tamkin, Alex, et al.
Published: (2024)
by: Tamkin, Alex, et al.
Published: (2024)
A Unifying Human-Centered AI Fairness Framework
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
A Conceptual Framework for Ethical Evaluation of Machine Learning Systems
by: Gupta, Neha R., et al.
Published: (2024)
by: Gupta, Neha R., et al.
Published: (2024)
AI Data Development: A Scorecard for the System Card Framework
by: Bahiru, Tadesse K., et al.
Published: (2025)
by: Bahiru, Tadesse K., et al.
Published: (2025)
A Post-Processing-Based Fair Federated Learning Framework
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
by: Farnadi, Golnoosh, et al.
Published: (2024)
by: Farnadi, Golnoosh, et al.
Published: (2024)
FairGridSearch: A Framework to Compare Fairness-Enhancing Models
by: Ma, Shih-Chi, et al.
Published: (2024)
by: Ma, Shih-Chi, et al.
Published: (2024)
Deprecating Benchmarks: Criteria and Framework
by: Joaquin, Ayrton San, et al.
Published: (2025)
by: Joaquin, Ayrton San, et al.
Published: (2025)
Equitable Evaluation via Elicitation
by: Du, Elbert, et al.
Published: (2026)
by: Du, Elbert, et al.
Published: (2026)
Evaluating Gemini in an arena for learning
by: LearnLM Team, et al.
Published: (2025)
by: LearnLM Team, et al.
Published: (2025)
Sabotage Evaluations for Frontier Models
by: Benton, Joe, et al.
Published: (2024)
by: Benton, Joe, et al.
Published: (2024)
Enhancing Team Diversity with Generative AI: A Novel Project Management Framework
by: Chan, Johnny, et al.
Published: (2025)
by: Chan, Johnny, et al.
Published: (2025)
Defining AI Models and AI Systems: A Framework to Resolve the Boundary Problem
by: Sun, Yuanyuan, et al.
Published: (2026)
by: Sun, Yuanyuan, et al.
Published: (2026)
A Comparative Study of Sampling Methods with Cross-Validation in the FedHome Framework
by: Ahmadi, Arash, et al.
Published: (2024)
by: Ahmadi, Arash, et al.
Published: (2024)
FlashST: A Simple and Universal Prompt-Tuning Framework for Traffic Prediction
by: Li, Zhonghang, et al.
Published: (2024)
by: Li, Zhonghang, et al.
Published: (2024)
Similar Items
-
What Does It Take to Build a Performant Selective Classifier?
by: Rabanser, Stephan, et al.
Published: (2025) -
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
by: Rabanser, Stephan
Published: (2025) -
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
by: Rabanser, Stephan, et al.
Published: (2025) -
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026) -
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
by: Patwardhan, Tejal, et al.
Published: (2025)