Deprecating Benchmarks: Criteria and Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Joaquin, Ayrton San, Gipiškis, Rokas, Staufer, Leon, Gil, Ariel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems
by: Gipiškis, Rokas, et al.
Published: (2024)
by: Gipiškis, Rokas, et al.
Published: (2024)
Evaluation Cards for XAI Metrics
by: Gipiškis, Rokas, et al.
Published: (2026)
by: Gipiškis, Rokas, et al.
Published: (2026)
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
by: Stelling, Lily, et al.
Published: (2025)
by: Stelling, Lily, et al.
Published: (2025)
Defining AI Models and AI Systems: A Framework to Resolve the Boundary Problem
by: Sun, Yuanyuan, et al.
Published: (2026)
by: Sun, Yuanyuan, et al.
Published: (2026)
The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act
by: Qureshi, Taro, et al.
Published: (2026)
by: Qureshi, Taro, et al.
Published: (2026)
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
by: Kelly, Christopher, et al.
Published: (2026)
by: Kelly, Christopher, et al.
Published: (2026)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Open Problems in Frontier AI Risk Management
by: Ziosi, Marta, et al.
Published: (2026)
by: Ziosi, Marta, et al.
Published: (2026)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models II: Benchmark Generation Process
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
by: Bae, Jaesung, et al.
Published: (2025)
by: Bae, Jaesung, et al.
Published: (2025)
A Comprehensive Survey and Classification of Evaluation Criteria for Trustworthy Artificial Intelligence
by: McCormack, Louise, et al.
Published: (2024)
by: McCormack, Louise, et al.
Published: (2024)
Algorithms for learning value-aligned policies considering admissibility relaxation
by: Holgado-Sánchez, Andrés, et al.
Published: (2024)
by: Holgado-Sánchez, Andrés, et al.
Published: (2024)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
by: Raman, Naveen, et al.
Published: (2026)
by: Raman, Naveen, et al.
Published: (2026)
DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
by: Cardei, Maria Ana, et al.
Published: (2025)
by: Cardei, Maria Ana, et al.
Published: (2025)
FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods
by: Han, Xiaotian, et al.
Published: (2023)
by: Han, Xiaotian, et al.
Published: (2023)
TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods
by: Qiu, Xiangfei, et al.
Published: (2024)
by: Qiu, Xiangfei, et al.
Published: (2024)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Scaling Legal AI: Benchmarking Mamba and Transformers for Statutory Classification and Case Law Retrieval
by: Maurya, Anuraj
Published: (2025)
by: Maurya, Anuraj
Published: (2025)
From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution
by: Koch, Bernard J., et al.
Published: (2024)
by: Koch, Bernard J., et al.
Published: (2024)
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
by: Ding, Shi, et al.
Published: (2025)
by: Ding, Shi, et al.
Published: (2025)
A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data
by: Tertulino, Rodrigo
Published: (2025)
by: Tertulino, Rodrigo
Published: (2025)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
by: Potham, Ram
Published: (2025)
by: Potham, Ram
Published: (2025)
A Unifying Human-Centered AI Fairness Framework
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
Automatic Evaluation Metrics for Artificially Generated Scientific Research
by: Höpner, Niklas, et al.
Published: (2025)
by: Höpner, Niklas, et al.
Published: (2025)
Designing Skill-Compatible AI: Methodologies and Frameworks in Chess
by: Hamade, Karim, et al.
Published: (2024)
by: Hamade, Karim, et al.
Published: (2024)
AI Data Development: A Scorecard for the System Card Framework
by: Bahiru, Tadesse K., et al.
Published: (2025)
by: Bahiru, Tadesse K., et al.
Published: (2025)
A Post-Processing-Based Fair Federated Learning Framework
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Say My Name: a Model's Bias Discovery Framework
by: Ciranni, Massimiliano, et al.
Published: (2024)
by: Ciranni, Massimiliano, et al.
Published: (2024)
Generative Artificial Intelligence: Evolving Technology, Growing Societal Impact, and Opportunities for Information Systems Research
by: Storey, Veda C., et al.
Published: (2025)
by: Storey, Veda C., et al.
Published: (2025)
AdvKT: An Adversarial Multi-Step Training Framework for Knowledge Tracing
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
FairGridSearch: A Framework to Compare Fairness-Enhancing Models
by: Ma, Shih-Chi, et al.
Published: (2024)
by: Ma, Shih-Chi, et al.
Published: (2024)
AI-powered Digital Framework for Personalized Economical Quality Learning at Scale
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
Enhancing Team Diversity with Generative AI: A Novel Project Management Framework
by: Chan, Johnny, et al.
Published: (2025)
by: Chan, Johnny, et al.
Published: (2025)
A Comparative Study of Sampling Methods with Cross-Validation in the FedHome Framework
by: Ahmadi, Arash, et al.
Published: (2024)
by: Ahmadi, Arash, et al.
Published: (2024)
FlashST: A Simple and Universal Prompt-Tuning Framework for Traffic Prediction
by: Li, Zhonghang, et al.
Published: (2024)
by: Li, Zhonghang, et al.
Published: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024)
by: Barrett, Anthony M., et al.
Published: (2024)
Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings
by: Pouget, Angéline, et al.
Published: (2025)
by: Pouget, Angéline, et al.
Published: (2025)
Exam Readiness Index (ERI): A Theoretical Framework for a Composite, Explainable Index
by: Verma, Ananda Prakash
Published: (2025)
by: Verma, Ananda Prakash
Published: (2025)
Similar Items
-
Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems
by: Gipiškis, Rokas, et al.
Published: (2024) -
Evaluation Cards for XAI Metrics
by: Gipiškis, Rokas, et al.
Published: (2026) -
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
by: Stelling, Lily, et al.
Published: (2025) -
Defining AI Models and AI Systems: A Framework to Resolve the Boundary Problem
by: Sun, Yuanyuan, et al.
Published: (2026) -
The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act
by: Qureshi, Taro, et al.
Published: (2026)