Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | McIntosh, Timothy R., Susnjak, Teo, Arachchilage, Nalin, Liu, Tong, Watters, Paul, Halgamuge, Malka N. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
by: McIntosh, Timothy R., et al.
Published: (2023)
by: McIntosh, Timothy R., et al.
Published: (2023)
From COBIT to ISO 42001: Evaluating Cybersecurity Frameworks for Opportunities, Risks, and Regulatory Compliance in Commercializing Large Language Models
by: McIntosh, Timothy R., et al.
Published: (2024)
by: McIntosh, Timothy R., et al.
Published: (2024)
A Design Science Blueprint for an Orchestrated AI Assistant in Doctoral Supervision
by: Susnjak, Teo, et al.
Published: (2025)
by: Susnjak, Teo, et al.
Published: (2025)
Why People Still Fall for Phishing Emails: An Empirical Investigation into How Users Make Email Response Decisions
by: Jayatilaka, Asangi, et al.
Published: (2024)
by: Jayatilaka, Asangi, et al.
Published: (2024)
Over the Edge of Chaos? Excess Complexity as a Roadblock to Artificial General Intelligence
by: Susnjak, Teo, et al.
Published: (2024)
by: Susnjak, Teo, et al.
Published: (2024)
The Good, the Bad, and the (Un)Usable: A Rapid Literature Review on Privacy as Code
by: Ferreyra, Nicolás E. Díaz, et al.
Published: (2024)
by: Ferreyra, Nicolás E. Díaz, et al.
Published: (2024)
Framework for Adoption of Generative Artificial Intelligence (GenAI) in Education
by: Shailendra, Samar, et al.
Published: (2024)
by: Shailendra, Samar, et al.
Published: (2024)
Bridging the Early Science Gap with Artificial Intelligence: Evaluating Large Language Models as Tools for Early Childhood Science Education
by: Bush, Annika, et al.
Published: (2025)
by: Bush, Annika, et al.
Published: (2025)
Situational Awareness as the Imperative Capability for Disaster Resilience in the Era of Complex Hazards and Artificial Intelligence
by: Pak, Hongrak, et al.
Published: (2025)
by: Pak, Hongrak, et al.
Published: (2025)
Making Data: The Work Behind Artificial Intelligence
by: Braz, Matheus Viana, et al.
Published: (2024)
by: Braz, Matheus Viana, et al.
Published: (2024)
GenAI Against Humanity: Nefarious Applications of Generative Artificial Intelligence and Large Language Models
by: Ferrara, Emilio
Published: (2023)
by: Ferrara, Emilio
Published: (2023)
Socially Minded Intelligence: How Individuals, Groups, and Artificial Intelligence Can Make Each Other Smarter (or Not)
by: Bingley, William J., et al.
Published: (2024)
by: Bingley, William J., et al.
Published: (2024)
"Can you be my mum?": Manipulating Social Robots in the Large Language Models Era
by: Abbo, Giulio Antonio, et al.
Published: (2025)
by: Abbo, Giulio Antonio, et al.
Published: (2025)
Artificial Intelligence Can Emulate Human Normative Judgments on Emotional Visual Scenes
by: Romeo, Zaira, et al.
Published: (2025)
by: Romeo, Zaira, et al.
Published: (2025)
The Role of Legal Frameworks in Shaping Ethical Artificial Intelligence Use in Corporate Governance
by: Mirishli, Shahmar
Published: (2025)
by: Mirishli, Shahmar
Published: (2025)
Artificial Intelligence in Spanish Gastroenterology: high expectations, limited integration. A national survey
by: Crespo, Javier, et al.
Published: (2026)
by: Crespo, Javier, et al.
Published: (2026)
Economic and Financial Learning with Artificial Intelligence: A Mixed-Methods Study on ChatGPT
by: Arndt, Holger
Published: (2024)
by: Arndt, Holger
Published: (2024)
Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing
by: Schmitt, Marc, et al.
Published: (2023)
by: Schmitt, Marc, et al.
Published: (2023)
Towards Synergistic Teacher-AI Interactions with Generative Artificial Intelligence
by: Cukurova, Mutlu, et al.
Published: (2025)
by: Cukurova, Mutlu, et al.
Published: (2025)
Tapping into the Natural Language System with Artificial Languages when Learning Programming
by: Hartmann, Elisa Madeleine, et al.
Published: (2024)
by: Hartmann, Elisa Madeleine, et al.
Published: (2024)
Critical Thinking in the Age of Artificial Intelligence: A Survey-Based Study with Machine Learning Insights
by: Bari, M Murshidul, et al.
Published: (2026)
by: Bari, M Murshidul, et al.
Published: (2026)
Artificial Intelligence in Environmental Protection: The Importance of Organizational Context from a Field Study in Wisconsin
by: Rothbacher, Nicolas, et al.
Published: (2025)
by: Rothbacher, Nicolas, et al.
Published: (2025)
Epistemological Fault Lines Between Human and Artificial Intelligence
by: Quattrociocchi, Walter, et al.
Published: (2025)
by: Quattrociocchi, Walter, et al.
Published: (2025)
Understanding and Evaluating Trust in Generative AI and Large Language Models for Spreadsheets
by: Thorne, Simon
Published: (2024)
by: Thorne, Simon
Published: (2024)
Cognitive Dissonance Artificial Intelligence (CD-AI): The Mind at War with Itself. Harnessing Discomfort to Sharpen Critical Thinking
by: Deliu, Delia
Published: (2025)
by: Deliu, Delia
Published: (2025)
Cultural Dimensions of Artificial Intelligence Adoption: Empirical Insights for Wave 1 from a Multinational Longitudinal Pilot Study
by: Cummings-Koether, Michelle J., et al.
Published: (2025)
by: Cummings-Koether, Michelle J., et al.
Published: (2025)
The Impact of Artificial Intelligence on Strategic Technology Management: A Mixed-Methods Analysis of Resources, Capabilities, and Human-AI Collaboration
by: Fascinari, Massimo, et al.
Published: (2025)
by: Fascinari, Massimo, et al.
Published: (2025)
GPTutor: Great Personalized Tutor with Large Language Models for Personalized Learning Content Generation
by: Chen, Eason, et al.
Published: (2024)
by: Chen, Eason, et al.
Published: (2024)
The Impact of Artificial Intelligence on Human Thought
by: Gesnot, Rénald
Published: (2025)
by: Gesnot, Rénald
Published: (2025)
Exploring Artificial Intelligence and Culture: Methodology for a comparative study of AI's impact on norms, trust, and problem-solving across academic and business environments
by: Huemmer, Matthias, et al.
Published: (2025)
by: Huemmer, Matthias, et al.
Published: (2025)
Evaluating Tenant-Landlord Tensions Using Generative AI on Online Tenant Forums
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
A Survey of Accessible Explainable Artificial Intelligence Research
by: Nwokoye, Chukwunonso Henry, et al.
Published: (2024)
by: Nwokoye, Chukwunonso Henry, et al.
Published: (2024)
Using Artificial Intelligence to Improve Classroom Learning Experience
by: Hossain, Shadeeb
Published: (2025)
by: Hossain, Shadeeb
Published: (2025)
False Sense of Security in Explainable Artificial Intelligence (XAI)
by: Chung, Neo Christopher, et al.
Published: (2024)
by: Chung, Neo Christopher, et al.
Published: (2024)
No General Code of Ethics for All: Ethical Considerations in Human-bot Psycho-counseling
by: Ma, Lizhi, et al.
Published: (2024)
by: Ma, Lizhi, et al.
Published: (2024)
Mapping Data Labour Supply Chain in Africa in an Era of Digital Apartheid: a Struggle for Recognition
by: Pidoux, Jessica, et al.
Published: (2025)
by: Pidoux, Jessica, et al.
Published: (2025)
The Ballad of the Bots: Sonification Using Cognitive Metaphor to Support Immersed Teleoperation of Robot Teams
by: Simmons, Joe, et al.
Published: (2024)
by: Simmons, Joe, et al.
Published: (2024)
Comuniqa : Exploring Large Language Models for improving speaking skills
by: Mhasakar, Manas, et al.
Published: (2024)
by: Mhasakar, Manas, et al.
Published: (2024)
Scaling Laws for Moral Machine Judgment in Large Language Models
by: Takemoto, Kazuhiro
Published: (2026)
by: Takemoto, Kazuhiro
Published: (2026)
Experiences from Integrating Large Language Model Chatbots into the Classroom
by: Hellas, Arto, et al.
Published: (2024)
by: Hellas, Arto, et al.
Published: (2024)
Similar Items
-
From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
by: McIntosh, Timothy R., et al.
Published: (2023) -
From COBIT to ISO 42001: Evaluating Cybersecurity Frameworks for Opportunities, Risks, and Regulatory Compliance in Commercializing Large Language Models
by: McIntosh, Timothy R., et al.
Published: (2024) -
A Design Science Blueprint for an Orchestrated AI Assistant in Doctoral Supervision
by: Susnjak, Teo, et al.
Published: (2025) -
Why People Still Fall for Phishing Emails: An Empirical Investigation into How Users Make Email Response Decisions
by: Jayatilaka, Asangi, et al.
Published: (2024) -
Over the Edge of Chaos? Excess Complexity as a Roadblock to Artificial General Intelligence
by: Susnjak, Teo, et al.
Published: (2024)