RAISE: A Unified Framework for Responsible AI Scoring and Evaluation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nguyen, Loc Phuc Truong, Do, Hung Thanh
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909861504090112
author Nguyen, Loc Phuc Truong
Do, Hung Thanh
author_facet Nguyen, Loc Phuc Truong
Do, Hung Thanh
contents As AI systems enter high-stakes domains, evaluation must extend beyond predictive accuracy to include explainability, fairness, robustness, and sustainability. We introduce RAISE (Responsible AI Scoring and Evaluation), a unified framework that quantifies model performance across these four dimensions and aggregates them into a single, holistic Responsibility Score. We evaluated three deep learning models: a Multilayer Perceptron (MLP), a Tabular ResNet, and a Feature Tokenizer Transformer, on structured datasets from finance, healthcare, and socioeconomics. Our findings reveal critical trade-offs: the MLP demonstrated strong sustainability and robustness, the Transformer excelled in explainability and fairness at a very high environmental cost, and the Tabular ResNet offered a balanced profile. These results underscore that no single model dominates across all responsibility criteria, highlighting the necessity of multi-dimensional evaluation for responsible model selection. Our implementation is available at: https://github.com/raise-framework/raise.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAISE: A Unified Framework for Responsible AI Scoring and Evaluation
Nguyen, Loc Phuc Truong
Do, Hung Thanh
Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computers and Society
As AI systems enter high-stakes domains, evaluation must extend beyond predictive accuracy to include explainability, fairness, robustness, and sustainability. We introduce RAISE (Responsible AI Scoring and Evaluation), a unified framework that quantifies model performance across these four dimensions and aggregates them into a single, holistic Responsibility Score. We evaluated three deep learning models: a Multilayer Perceptron (MLP), a Tabular ResNet, and a Feature Tokenizer Transformer, on structured datasets from finance, healthcare, and socioeconomics. Our findings reveal critical trade-offs: the MLP demonstrated strong sustainability and robustness, the Transformer excelled in explainability and fairness at a very high environmental cost, and the Tabular ResNet offered a balanced profile. These results underscore that no single model dominates across all responsibility criteria, highlighting the necessity of multi-dimensional evaluation for responsible model selection. Our implementation is available at: https://github.com/raise-framework/raise.
title RAISE: A Unified Framework for Responsible AI Scoring and Evaluation
topic Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computers and Society
url https://arxiv.org/abs/2510.18559