AI Act Evaluation Benchmark: An Open, Transparent, and Reproducible Evaluation Dataset for NLP and RAG Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Davvetas, Athanasios, Papademas, Michael, Ziouvelou, Xenia, Karkaletsis, Vangelis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Ethical Principles and Algorithmic Methods: An Alternative Approach for Assessing Trustworthiness in AI Systems
by: Papademas, Michael, et al.
Published: (2025)
by: Papademas, Michael, et al.
Published: (2025)
TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment
by: Davvetas, Athanasios, et al.
Published: (2025)
by: Davvetas, Athanasios, et al.
Published: (2025)
A Systematic Evaluation of LLM Strategies for Mental Health Text Analysis: Fine-tuning vs. Prompt Engineering vs. RAG
by: Kermani, Arshia, et al.
Published: (2025)
by: Kermani, Arshia, et al.
Published: (2025)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
by: Srivastav, Vaibhav, et al.
Published: (2025)
by: Srivastav, Vaibhav, et al.
Published: (2025)
AI, Climate, and Transparency: Operationalizing and Improving the AI Act
by: Alder, Nicolas, et al.
Published: (2024)
by: Alder, Nicolas, et al.
Published: (2024)
AI Benchmarks and Datasets for LLM Evaluation
by: Ivanov, Todor, et al.
Published: (2024)
by: Ivanov, Todor, et al.
Published: (2024)
FeDa4Fair: Client-Level Federated Datasets for Fairness Evaluation
by: Heilmann, Xenia, et al.
Published: (2025)
by: Heilmann, Xenia, et al.
Published: (2025)
Benchmark Transparency: Measuring the Impact of Data on Evaluation
by: Kovatchev, Venelin, et al.
Published: (2024)
by: Kovatchev, Venelin, et al.
Published: (2024)
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
by: Brehme, Lorenz, et al.
Published: (2025)
by: Brehme, Lorenz, et al.
Published: (2025)
Ensuring Reproducibility in Generative AI Systems for General Use Cases: A Framework for Regression Testing and Open Datasets
by: Morishige, Masumi, et al.
Published: (2025)
by: Morishige, Masumi, et al.
Published: (2025)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
by: Liu, Jiayi, et al.
Published: (2026)
by: Liu, Jiayi, et al.
Published: (2026)
Evaluating a Methodology for Increasing AI Transparency: A Case Study
by: Piorkowski, David, et al.
Published: (2022)
by: Piorkowski, David, et al.
Published: (2022)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
by: Agarwal, Chirag, et al.
Published: (2022)
by: Agarwal, Chirag, et al.
Published: (2022)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
by: Li, Jingyue, et al.
Published: (2026)
by: Li, Jingyue, et al.
Published: (2026)
StratRAG: A Multi-Hop Retrieval Evaluation Dataset for Retrieval-Augmented Generation Systems
by: Patodiya, Aryan
Published: (2026)
by: Patodiya, Aryan
Published: (2026)
GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation
by: Sorodoc, Ionut-Teodor, et al.
Published: (2025)
by: Sorodoc, Ionut-Teodor, et al.
Published: (2025)
AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
by: Wu, Zonghan, et al.
Published: (2025)
by: Wu, Zonghan, et al.
Published: (2025)
The Model Openness Framework: Promoting Completeness and Openness for Reproducibility, Transparency, and Usability in Artificial Intelligence
by: White, Matt, et al.
Published: (2024)
by: White, Matt, et al.
Published: (2024)
Responsible AI in NLP: GUS-Net Span-Level Bias Detection Dataset and Benchmark for Generalizations, Unfairness, and Stereotypes
by: Powers, Maximus, et al.
Published: (2024)
by: Powers, Maximus, et al.
Published: (2024)
UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
Framing AI System Benchmarking as a Learning Task: FlexBench and the Open MLPerf Dataset
by: Fursin, Grigori, et al.
Published: (2025)
by: Fursin, Grigori, et al.
Published: (2025)
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
by: An, Subin, et al.
Published: (2025)
by: An, Subin, et al.
Published: (2025)
Knowledge-Graph Based RAG System Evaluation Framework
by: Dong, Sicheng, et al.
Published: (2025)
by: Dong, Sicheng, et al.
Published: (2025)
Evaluating Saliency Explanations in NLP by Crowdsourcing
by: Lu, Xiaotian, et al.
Published: (2024)
by: Lu, Xiaotian, et al.
Published: (2024)
RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
by: Zhang, Rongyang, et al.
Published: (2025)
by: Zhang, Rongyang, et al.
Published: (2025)
BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation
by: Metcalf, Sara, et al.
Published: (2026)
by: Metcalf, Sara, et al.
Published: (2026)
Procrustean Bed for AI-Driven Retrosynthesis: A Unified Framework for Reproducible Evaluation
by: Morgunov, Anton, et al.
Published: (2025)
by: Morgunov, Anton, et al.
Published: (2025)
Select, Label, Evaluate: Active Testing in NLP
by: Purificato, Antonio, et al.
Published: (2026)
by: Purificato, Antonio, et al.
Published: (2026)
Evaluation Metrics for Text Data Augmentation in NLP
by: Amadeus, Marcellus, et al.
Published: (2024)
by: Amadeus, Marcellus, et al.
Published: (2024)
Human-Centered Evaluation of RAG outputs: a framework and questionnaire for human-AI collaboration
by: Mangold, Aline, et al.
Published: (2025)
by: Mangold, Aline, et al.
Published: (2025)
Toward Trustworthy Evaluation of Sustainability Rating Methodologies: A Human-AI Collaborative Framework for Benchmark Dataset Construction
by: Cai, Xiaoran, et al.
Published: (2026)
by: Cai, Xiaoran, et al.
Published: (2026)
SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Open-World Evaluations for Measuring Frontier AI Capabilities
by: Kapoor, Sayash, et al.
Published: (2026)
by: Kapoor, Sayash, et al.
Published: (2026)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
by: Hu, Junyi, et al.
Published: (2024)
by: Hu, Junyi, et al.
Published: (2024)
CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning
by: Panayiotou, Panayiotis, et al.
Published: (2025)
by: Panayiotou, Panayiotis, et al.
Published: (2025)
Deepchecks: Evaluating Retrieval-Augmented Generation (RAG)
by: Gerner, Assaf, et al.
Published: (2026)
by: Gerner, Assaf, et al.
Published: (2026)
RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition
by: Cofala, Tim, et al.
Published: (2025)
by: Cofala, Tim, et al.
Published: (2025)
Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Similar Items
-
Bridging Ethical Principles and Algorithmic Methods: An Alternative Approach for Assessing Trustworthiness in AI Systems
by: Papademas, Michael, et al.
Published: (2025) -
TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment
by: Davvetas, Athanasios, et al.
Published: (2025) -
A Systematic Evaluation of LLM Strategies for Mental Health Text Analysis: Fine-tuning vs. Prompt Engineering vs. RAG
by: Kermani, Arshia, et al.
Published: (2025) -
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
by: Srivastav, Vaibhav, et al.
Published: (2025) -
AI, Climate, and Transparency: Operationalizing and Improving the AI Act
by: Alder, Nicolas, et al.
Published: (2024)