Exposing Assumptions in AI Benchmarks through Cognitive Modelling
Fuente:
arXiv
Saved in:
| Main Authors: | Rystrøm, Jonathan H., Enevoldsen, Kenneth C. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounding Text Embeddings in Stakeholder Associations
by: Rystrøm, Jonathan, et al.
Published: (2026)
by: Rystrøm, Jonathan, et al.
Published: (2026)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
by: Enevoldsen, Kenneth, et al.
Published: (2024)
by: Enevoldsen, Kenneth, et al.
Published: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
by: Chung, Isaac, et al.
Published: (2025)
by: Chung, Isaac, et al.
Published: (2025)
Oversight Structures for Agentic AI in Public-Sector Organizations
by: Schmitz, Chris, et al.
Published: (2025)
by: Schmitz, Chris, et al.
Published: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
by: Nielsen, Dan Saattrup, et al.
Published: (2024)
by: Nielsen, Dan Saattrup, et al.
Published: (2024)
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning
by: Wang, Olivia Peiyu, et al.
Published: (2026)
by: Wang, Olivia Peiyu, et al.
Published: (2026)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
Agent Benchmarks Fail Public Sector Requirements
by: Rystrøm, Jonathan, et al.
Published: (2026)
by: Rystrøm, Jonathan, et al.
Published: (2026)
Cognitive Effects in Large Language Models
by: Shaki, Jonathan, et al.
Published: (2023)
by: Shaki, Jonathan, et al.
Published: (2023)
Cognitive LLMs: Towards Integrating Cognitive Architectures and Large Language Models for Manufacturing Decision-making
by: Wu, Siyu, et al.
Published: (2024)
by: Wu, Siyu, et al.
Published: (2024)
Hallucination is Inevitable for LLMs with the Open World Assumption
by: Xu, Bowen
Published: (2025)
by: Xu, Bowen
Published: (2025)
BrainBench: Exposing the Commonsense Reasoning Gap in Large Language Models
by: Tang, Yuzhe
Published: (2026)
by: Tang, Yuzhe
Published: (2026)
An analysis of AI Decision under Risk: Prospect theory emerges in Large Language Models
by: Payne, Kenneth
Published: (2025)
by: Payne, Kenneth
Published: (2025)
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation
by: Zheng, Huawei, et al.
Published: (2026)
by: Zheng, Huawei, et al.
Published: (2026)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
by: Dong, Nguyen Tien, et al.
Published: (2025)
by: Dong, Nguyen Tien, et al.
Published: (2025)
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
by: Bean, Andrew M., et al.
Published: (2025)
by: Bean, Andrew M., et al.
Published: (2025)
EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education
by: Ma, Guoqing, et al.
Published: (2025)
by: Ma, Guoqing, et al.
Published: (2025)
A Question on the Explainability of Large Language Models and the Word-Level Univariate First-Order Plausibility Assumption
by: Bogaert, Jeremie, et al.
Published: (2024)
by: Bogaert, Jeremie, et al.
Published: (2024)
How Small Transformation Expose the Weakness of Semantic Similarity Measures
by: Nikiema, Serge Lionel, et al.
Published: (2025)
by: Nikiema, Serge Lionel, et al.
Published: (2025)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
by: Chang, Chen-Chi, et al.
Published: (2024)
by: Chang, Chen-Chi, et al.
Published: (2024)
The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows
by: Kim, Hyunwoo, et al.
Published: (2026)
by: Kim, Hyunwoo, et al.
Published: (2026)
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
by: Mor-Lan, Guy, et al.
Published: (2026)
by: Mor-Lan, Guy, et al.
Published: (2026)
HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning
by: Yang, Qihao, et al.
Published: (2025)
by: Yang, Qihao, et al.
Published: (2025)
TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
by: Wu, Siyang, et al.
Published: (2025)
by: Wu, Siyang, et al.
Published: (2025)
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
by: Zhao, Boxiang, et al.
Published: (2026)
by: Zhao, Boxiang, et al.
Published: (2026)
Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias
by: Ke, Yu He, et al.
Published: (2024)
by: Ke, Yu He, et al.
Published: (2024)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
by: Furuta, Hiroki, et al.
Published: (2023)
by: Furuta, Hiroki, et al.
Published: (2023)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment
by: Yaacoub, Antoun, et al.
Published: (2025)
by: Yaacoub, Antoun, et al.
Published: (2025)
How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games
by: Xie, Yutong, et al.
Published: (2024)
by: Xie, Yutong, et al.
Published: (2024)
MAEB: Massive Audio Embedding Benchmark
by: Assadi, Adnan El, et al.
Published: (2026)
by: Assadi, Adnan El, et al.
Published: (2026)
M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark
by: Song, Wei, et al.
Published: (2024)
by: Song, Wei, et al.
Published: (2024)
INTIMA: A Benchmark for Human-AI Companionship Behavior
by: Kaffee, Lucie-Aimée, et al.
Published: (2025)
by: Kaffee, Lucie-Aimée, et al.
Published: (2025)
Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations
by: Dimgba, Martha O., et al.
Published: (2025)
by: Dimgba, Martha O., et al.
Published: (2025)
Similar Items
-
Grounding Text Embeddings in Stakeholder Associations
by: Rystrøm, Jonathan, et al.
Published: (2026) -
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
by: Enevoldsen, Kenneth, et al.
Published: (2024) -
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025) -
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
by: Chung, Isaac, et al.
Published: (2025) -
Oversight Structures for Agentic AI in Public-Sector Organizations
by: Schmitz, Chris, et al.
Published: (2025)