Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Salaudeen, Olawale, Reuel, Anka, Ahmed, Ahmed, Bedi, Suhana, Robertson, Zachary, Sundar, Sudharsan, Domingue, Ben, Wang, Angelina, Koyejo, Sanmi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Causally Inspired Regularization Enables Domain General Representations
por: Salaudeen, Olawale, et al.
Publicado: (2024)
por: Salaudeen, Olawale, et al.
Publicado: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
por: Robertson, Zachary, et al.
Publicado: (2025)
por: Robertson, Zachary, et al.
Publicado: (2025)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
por: Salaudeen, Olawale, et al.
Publicado: (2025)
por: Salaudeen, Olawale, et al.
Publicado: (2025)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
por: Hardy, Michael, et al.
Publicado: (2026)
por: Hardy, Michael, et al.
Publicado: (2026)
Fairness in Reinforcement Learning: A Survey
por: Reuel, Anka, et al.
Publicado: (2024)
por: Reuel, Anka, et al.
Publicado: (2024)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
por: Haupt, Andreas, et al.
Publicado: (2026)
por: Haupt, Andreas, et al.
Publicado: (2026)
Generative AI Needs Adaptive Governance
por: Reuel, Anka, et al.
Publicado: (2024)
por: Reuel, Anka, et al.
Publicado: (2024)
Toward an Evaluation Science for Generative AI Systems
por: Weidinger, Laura, et al.
Publicado: (2025)
por: Weidinger, Laura, et al.
Publicado: (2025)
Fantastic Bugs and Where to Find Them in AI Benchmarks
por: Truong, Sang, et al.
Publicado: (2025)
por: Truong, Sang, et al.
Publicado: (2025)
Audit Cards: Contextualizing AI Evaluations
por: Staufer, Leon, et al.
Publicado: (2025)
por: Staufer, Leon, et al.
Publicado: (2025)
Position: Model Collapse Does Not Mean What You Think
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
por: Reuel, Anka, et al.
Publicado: (2024)
por: Reuel, Anka, et al.
Publicado: (2024)
What's in a Query: Polarity-Aware Distribution-Based Fair Ranking
por: Balagopalan, Aparna, et al.
Publicado: (2025)
por: Balagopalan, Aparna, et al.
Publicado: (2025)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
por: Wang, Angelina, et al.
Publicado: (2025)
por: Wang, Angelina, et al.
Publicado: (2025)
The Optimization Paradox in Clinical AI Multi-Agent Systems
por: Bedi, Suhana, et al.
Publicado: (2025)
por: Bedi, Suhana, et al.
Publicado: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
por: Obbad, Elyas, et al.
Publicado: (2024)
por: Obbad, Elyas, et al.
Publicado: (2024)
Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
por: Pfohl, Stephen R., et al.
Publicado: (2025)
por: Pfohl, Stephen R., et al.
Publicado: (2025)
ImageNot: A contrast with ImageNet preserves model rankings
por: Salaudeen, Olawale, et al.
Publicado: (2024)
por: Salaudeen, Olawale, et al.
Publicado: (2024)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
por: Lamparth, Max, et al.
Publicado: (2023)
por: Lamparth, Max, et al.
Publicado: (2023)
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
por: Reuel, Anka, et al.
Publicado: (2025)
por: Reuel, Anka, et al.
Publicado: (2025)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
por: Miranda, Brando, et al.
Publicado: (2023)
por: Miranda, Brando, et al.
Publicado: (2023)
Proxy Methods for Domain Adaptation
por: Tsai, Katherine, et al.
Publicado: (2024)
por: Tsai, Katherine, et al.
Publicado: (2024)
Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHRs
por: Wornow, Michael, et al.
Publicado: (2024)
por: Wornow, Michael, et al.
Publicado: (2024)
On Fairness of Low-Rank Adaptation of Large Models
por: Ding, Zhoujie, et al.
Publicado: (2024)
por: Ding, Zhoujie, et al.
Publicado: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
por: Chawla, Krrish, et al.
Publicado: (2025)
por: Chawla, Krrish, et al.
Publicado: (2025)
Extracting books from production language models
por: Ahmed, Ahmed, et al.
Publicado: (2026)
por: Ahmed, Ahmed, et al.
Publicado: (2026)
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
por: Sühr, Tom, et al.
Publicado: (2025)
por: Sühr, Tom, et al.
Publicado: (2025)
Implicit Regularization in Feedback Alignment Learning Mechanisms for Neural Networks
por: Robertson, Zachary, et al.
Publicado: (2023)
por: Robertson, Zachary, et al.
Publicado: (2023)
Globalizing Fairness Attributes in Machine Learning: A Case Study on Health in Africa
por: Asiedu, Mercy Nyamewaa, et al.
Publicado: (2023)
por: Asiedu, Mercy Nyamewaa, et al.
Publicado: (2023)
Responsible AI in the Global Context: Maturity Model and Survey
por: Reuel, Anka, et al.
Publicado: (2024)
por: Reuel, Anka, et al.
Publicado: (2024)
Shaping AI's Impact on Billions of Lives
por: Cuéllar, Mariano-Florentino, et al.
Publicado: (2024)
por: Cuéllar, Mariano-Florentino, et al.
Publicado: (2024)
Scalable Ensembling For Mitigating Reward Overoptimisation
por: Ahmed, Ahmed M., et al.
Publicado: (2024)
por: Ahmed, Ahmed M., et al.
Publicado: (2024)
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
por: Meimandi, Kiana Jafari, et al.
Publicado: (2025)
por: Meimandi, Kiana Jafari, et al.
Publicado: (2025)
A Framework for Objective-Driven Dynamical Stochastic Fields
por: Zhang, Yibo Jacky, et al.
Publicado: (2025)
por: Zhang, Yibo Jacky, et al.
Publicado: (2025)
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
por: Tang, Zeyu, et al.
Publicado: (2025)
por: Tang, Zeyu, et al.
Publicado: (2025)
Semantic Risk Scoring of Aggregated Metrics: An AI-Driven Approach for Healthcare Data Governance
por: Ahmed, Mohammed Omer Shakeel
Publicado: (2026)
por: Ahmed, Mohammed Omer Shakeel
Publicado: (2026)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
por: Vo, Truong, et al.
Publicado: (2025)
por: Vo, Truong, et al.
Publicado: (2025)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
por: Wang, Angelina, et al.
Publicado: (2025)
por: Wang, Angelina, et al.
Publicado: (2025)
Pretraining Scaling Laws for Generative Evaluations of Language Models
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
AI Evaluation Should Require Standardized Item-Level Data Releases
por: Jiang, Han, et al.
Publicado: (2026)
por: Jiang, Han, et al.
Publicado: (2026)
Ejemplares similares
-
Causally Inspired Regularization Enables Domain General Representations
por: Salaudeen, Olawale, et al.
Publicado: (2024) -
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
por: Robertson, Zachary, et al.
Publicado: (2025) -
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
por: Salaudeen, Olawale, et al.
Publicado: (2025) -
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
por: Hardy, Michael, et al.
Publicado: (2026) -
Fairness in Reinforcement Learning: A Survey
por: Reuel, Anka, et al.
Publicado: (2024)