Evaluating AI Evaluation: Perils and Prospects
Fuente:
arXiv
Saved in:
| Main Author: | Burden, John |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education
by: Grume, Juvy C., et al.
Published: (2026)
by: Grume, Juvy C., et al.
Published: (2026)
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
by: Jalilian, Laleh, et al.
Published: (2024)
by: Jalilian, Laleh, et al.
Published: (2024)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
Evaluating a Methodology for Increasing AI Transparency: A Case Study
by: Piorkowski, David, et al.
Published: (2022)
by: Piorkowski, David, et al.
Published: (2022)
No Size Fits All: The Perils and Pitfalls of Leveraging LLMs Vary with Company Size
by: Urlana, Ashok, et al.
Published: (2024)
by: Urlana, Ashok, et al.
Published: (2024)
A Knowledge-Component-Based Methodology for Evaluating AI Assistants
by: Qi, Laryn, et al.
Published: (2024)
by: Qi, Laryn, et al.
Published: (2024)
Evaluating General-Purpose AI with Psychometrics
by: Wang, Xiting, et al.
Published: (2023)
by: Wang, Xiting, et al.
Published: (2023)
Responsible Evaluation of AI for Mental Health
by: Arnaout, Hiba, et al.
Published: (2026)
by: Arnaout, Hiba, et al.
Published: (2026)
Potential and Perils of Large Language Models as Judges of Unstructured Textual Data
by: Bedemariam, Rewina, et al.
Published: (2025)
by: Bedemariam, Rewina, et al.
Published: (2025)
Justifications for Democratizing AI Alignment and Their Prospects
by: Steingrüber, André, et al.
Published: (2025)
by: Steingrüber, André, et al.
Published: (2025)
Comprehensive Framework for Evaluating Conversational AI Chatbots
by: Gupta, Shailja, et al.
Published: (2025)
by: Gupta, Shailja, et al.
Published: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
by: Costa, Mariana Lins
Published: (2026)
by: Costa, Mariana Lins
Published: (2026)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
Exploring AI Writers: Technology, Impact, and Future Prospects
by: Huang, Zhiqian
Published: (2025)
by: Huang, Zhiqian
Published: (2025)
Evaluation Framework for AI Systems in "the Wild"
by: Jabbour, Sarah, et al.
Published: (2025)
by: Jabbour, Sarah, et al.
Published: (2025)
Evaluating the Social Impact of Generative AI Systems in Systems and Society
by: Solaiman, Irene, et al.
Published: (2023)
by: Solaiman, Irene, et al.
Published: (2023)
A Study on the Framework for Evaluating the Ethics and Trustworthiness of Generative AI
by: Jeong, Cheonsu, et al.
Published: (2025)
by: Jeong, Cheonsu, et al.
Published: (2025)
Computational Hermeneutics: Evaluating generative AI as a cultural technology
by: Kommers, Cody, et al.
Published: (2026)
by: Kommers, Cody, et al.
Published: (2026)
Do Generative AI Tools Ensure Green Code? An Investigative Study
by: Sikand, Samarth, et al.
Published: (2025)
by: Sikand, Samarth, et al.
Published: (2025)
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources
by: Clark, Hannah-Beth, et al.
Published: (2025)
by: Clark, Hannah-Beth, et al.
Published: (2025)
Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI
by: O'Herlihy, Michael, et al.
Published: (2026)
by: O'Herlihy, Michael, et al.
Published: (2026)
Frontier AI Ethics: Anticipating and Evaluating the Societal Impacts of Language Model Agents
by: Lazar, Seth
Published: (2024)
by: Lazar, Seth
Published: (2024)
Human Experts' Evaluation of Generative AI for Contextualizing STEAM Education in the Global South
by: Nyaaba, Matthew, et al.
Published: (2025)
by: Nyaaba, Matthew, et al.
Published: (2025)
AI-Based Reconstruction from Inherited Personal Data: Analysis, Feasibility, and Prospects
by: Zilberman, Mark
Published: (2025)
by: Zilberman, Mark
Published: (2025)
Dialogue with the Machine and Dialogue with the Art World: Evaluating Generative AI for Culturally-Situated Creativity
by: Qadri, Rida, et al.
Published: (2024)
by: Qadri, Rida, et al.
Published: (2024)
A Framework for Human-AI Q-Matrix Refinement: A NeuralCDM Evaluation
by: Zhang, Ying, et al.
Published: (2026)
by: Zhang, Ying, et al.
Published: (2026)
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
by: McCaslin, Tegan, et al.
Published: (2025)
by: McCaslin, Tegan, et al.
Published: (2025)
Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education
by: Chamberland, Jean-Francois, et al.
Published: (2025)
by: Chamberland, Jean-Francois, et al.
Published: (2025)
The Case for "Thick Evaluations" of Cultural Representation in AI
by: Qadri, Rida, et al.
Published: (2025)
by: Qadri, Rida, et al.
Published: (2025)
Evaluating undergraduate mathematics examinations in the era of generative AI: a curriculum-level case study
by: Walker, Benjamin J., et al.
Published: (2025)
by: Walker, Benjamin J., et al.
Published: (2025)
LegalScore: Development of a Benchmark for Evaluating AI Models in Legal Career Exams in Brazil
by: Caparroz, Roberto, et al.
Published: (2025)
by: Caparroz, Roberto, et al.
Published: (2025)
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
by: Schwartz, Reva, et al.
Published: (2025)
by: Schwartz, Reva, et al.
Published: (2025)
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
by: Choong, Yee-Yin, et al.
Published: (2026)
by: Choong, Yee-Yin, et al.
Published: (2026)
AI Evaluation Should Require Standardized Item-Level Data Releases
by: Jiang, Han, et al.
Published: (2026)
by: Jiang, Han, et al.
Published: (2026)
Critically Engaged Pragmatism: A Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools
by: Lee, Carole J.
Published: (2026)
by: Lee, Carole J.
Published: (2026)
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma
by: Schwartz, Reva, et al.
Published: (2026)
by: Schwartz, Reva, et al.
Published: (2026)
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents
by: Vatsal, Shubham, et al.
Published: (2026)
by: Vatsal, Shubham, et al.
Published: (2026)
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
by: Reuel, Anka, et al.
Published: (2025)
by: Reuel, Anka, et al.
Published: (2025)
Unveiling User Perceptions in the Generative AI Era: A Sentiment-Driven Evaluation of AI Educational Apps' Role in Digital Transformation of e-Teaching
by: Mazaherian, Adeleh, et al.
Published: (2025)
by: Mazaherian, Adeleh, et al.
Published: (2025)
Similar Items
-
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
by: Petrova, Nora, et al.
Published: (2026) -
Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education
by: Grume, Juvy C., et al.
Published: (2026) -
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
by: Jalilian, Laleh, et al.
Published: (2024) -
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025) -
Evaluating a Methodology for Increasing AI Transparency: A Case Study
by: Piorkowski, David, et al.
Published: (2022)