Automatic Evaluation Metrics for Artificially Generated Scientific Research
Fuente:
arXiv
Saved in:
| Main Authors: | Höpner, Niklas, Eshuijs, Leon, Alivanistos, Dimitrios, Zamprogno, Giacomo, Tiddi, Ilaria |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Augmentation for Instruction Following Policies via Trajectory Segmentation
by: Höpner, Niklas, et al.
Published: (2025)
by: Höpner, Niklas, et al.
Published: (2025)
Generative Artificial Intelligence: Evolving Technology, Growing Societal Impact, and Opportunities for Information Systems Research
by: Storey, Veda C., et al.
Published: (2025)
by: Storey, Veda C., et al.
Published: (2025)
Balancing the Scales: Reinforcement Learning for Fair Classification
by: Eshuijs, Leon, et al.
Published: (2024)
by: Eshuijs, Leon, et al.
Published: (2024)
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
by: Wang, Miles, et al.
Published: (2026)
by: Wang, Miles, et al.
Published: (2026)
SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
by: Yu, Ke, et al.
Published: (2025)
by: Yu, Ke, et al.
Published: (2025)
Generative Artificial Intelligence in Healthcare: Ethical Considerations and Assessment Checklist
by: Ning, Yilin, et al.
Published: (2023)
by: Ning, Yilin, et al.
Published: (2023)
Deep Learning Opacity in Scientific Discovery
by: Duede, Eamon
Published: (2022)
by: Duede, Eamon
Published: (2022)
Evaluation Cards for XAI Metrics
by: Gipiškis, Rokas, et al.
Published: (2026)
by: Gipiškis, Rokas, et al.
Published: (2026)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups
by: Popoola, Gideon, et al.
Published: (2026)
by: Popoola, Gideon, et al.
Published: (2026)
The Mirage of Artificial Intelligence Terms of Use Restrictions
by: Henderson, Peter, et al.
Published: (2024)
by: Henderson, Peter, et al.
Published: (2024)
Automatically Inferring Teachers' Geometric Content Knowledge: A Skills Based Approach
by: Fenigstein, Ziv, et al.
Published: (2026)
by: Fenigstein, Ziv, et al.
Published: (2026)
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
by: Jurenka, Irina, et al.
Published: (2024)
by: Jurenka, Irina, et al.
Published: (2024)
A Trustworthiness-based Metaphysics of Artificial Intelligence Systems
by: Ferrario, Andrea
Published: (2025)
by: Ferrario, Andrea
Published: (2025)
Artificial Intelligence Ecosystem for Automating Self-Directed Teaching
by: Gotavade, Tejas Satish
Published: (2024)
by: Gotavade, Tejas Satish
Published: (2024)
Responsible Artificial Intelligence: A Structured Literature Review
by: Goellner, Sabrina, et al.
Published: (2024)
by: Goellner, Sabrina, et al.
Published: (2024)
The Pursuit of Fairness in Artificial Intelligence Models: A Survey
by: Kheya, Tahsin Alamgir, et al.
Published: (2024)
by: Kheya, Tahsin Alamgir, et al.
Published: (2024)
A Comprehensive Survey and Classification of Evaluation Criteria for Trustworthy Artificial Intelligence
by: McCormack, Louise, et al.
Published: (2024)
by: McCormack, Louise, et al.
Published: (2024)
Artificial intelligence and democracy: Towards digital authoritarianism or a democratic upgrade?
by: Panagopoulou, Fereniki
Published: (2025)
by: Panagopoulou, Fereniki
Published: (2025)
Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents
by: Baum, Kevin, et al.
Published: (2024)
by: Baum, Kevin, et al.
Published: (2024)
New-Onset Diabetes Assessment Using Artificial Intelligence-Enhanced Electrocardiography
by: Zhang, Hao, et al.
Published: (2022)
by: Zhang, Hao, et al.
Published: (2022)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Evaluating Retrieval-Augmented Generation Strategies for Large Language Models in Travel Mode Choice Prediction
by: Xu, Yiming, et al.
Published: (2025)
by: Xu, Yiming, et al.
Published: (2025)
Mapping the Potential of Explainable AI for Fairness Along the AI Lifecycle
by: Deck, Luca, et al.
Published: (2024)
by: Deck, Luca, et al.
Published: (2024)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
Backdoor for Debias: Mitigating Model Bias with Backdoor Attack-based Artificial Bias
by: Wu, Shangxi, et al.
Published: (2023)
by: Wu, Shangxi, et al.
Published: (2023)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026)
by: Bai, Xiaoyan, et al.
Published: (2026)
Vision Paper: Designing Graph Neural Networks in Compliance with the European Artificial Intelligence Act
by: Hoffmann, Barbara, et al.
Published: (2024)
by: Hoffmann, Barbara, et al.
Published: (2024)
Developing and Deploying Industry Standards for Artificial Intelligence in Education (AIED): Challenges, Strategies, and Future Directions
by: Tong, Richard, et al.
Published: (2024)
by: Tong, Richard, et al.
Published: (2024)
Mind the Gap! Bridging Explainable Artificial Intelligence and Human Understanding with Luhmann's Functional Theory of Communication
by: Keenan, Bernard, et al.
Published: (2023)
by: Keenan, Bernard, et al.
Published: (2023)
Evaluating the Generalization Ability of Spatiotemporal Model in Urban Scenario
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
Governance of Generative Artificial Intelligence for Companies
by: Schneider, Johannes, et al.
Published: (2024)
by: Schneider, Johannes, et al.
Published: (2024)
Can Large Language Models Unlock Novel Scientific Research Ideas?
by: Kumar, Sandeep, et al.
Published: (2024)
by: Kumar, Sandeep, et al.
Published: (2024)
Generative AI in Health Economics and Outcomes Research: A Taxonomy of Key Definitions and Emerging Applications, an ISPOR Working Group Report
by: Fleurence, Rachael, et al.
Published: (2024)
by: Fleurence, Rachael, et al.
Published: (2024)
What is Reproducibility in Artificial Intelligence and Machine Learning Research?
by: Desai, Abhyuday, et al.
Published: (2024)
by: Desai, Abhyuday, et al.
Published: (2024)
FairMT: Fairness for Heterogeneous Multi-Task Learning
by: Hu, Guanyu, et al.
Published: (2025)
by: Hu, Guanyu, et al.
Published: (2025)
The Quest for Reliable Metrics of Responsible AI
by: Rampisela, Theresia Veronika, et al.
Published: (2025)
by: Rampisela, Theresia Veronika, et al.
Published: (2025)
Evaluating Gemini in an arena for learning
by: LearnLM Team, et al.
Published: (2025)
by: LearnLM Team, et al.
Published: (2025)
Similar Items
-
Data Augmentation for Instruction Following Policies via Trajectory Segmentation
by: Höpner, Niklas, et al.
Published: (2025) -
Generative Artificial Intelligence: Evolving Technology, Growing Societal Impact, and Opportunities for Information Systems Research
by: Storey, Veda C., et al.
Published: (2025) -
Balancing the Scales: Reinforcement Learning for Fair Classification
by: Eshuijs, Leon, et al.
Published: (2024) -
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
by: Wang, Miles, et al.
Published: (2026) -
SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
by: Yu, Ke, et al.
Published: (2025)