A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Haznitrama, Faiz Ghifari, Ardi, Faeyza Rishad, Oh, Alice |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OLA: Output Language Alignment in Code-Switched LLM Interactions
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
von: Oh, Juhyun, et al.
Veröffentlicht: (2026)
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese
von: Putri, Rifki Afina, et al.
Veröffentlicht: (2024)
von: Putri, Rifki Afina, et al.
Veröffentlicht: (2024)
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
von: Li, Zheng, et al.
Veröffentlicht: (2026)
von: Li, Zheng, et al.
Veröffentlicht: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
von: Valentin, Thomas, et al.
Veröffentlicht: (2025)
von: Valentin, Thomas, et al.
Veröffentlicht: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
von: Jin, Jiho, et al.
Veröffentlicht: (2026)
von: Jin, Jiho, et al.
Veröffentlicht: (2026)
Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
von: Wu, Yunze, et al.
Veröffentlicht: (2025)
von: Wu, Yunze, et al.
Veröffentlicht: (2025)
The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective
von: Imran, Muhammad Osama, et al.
Veröffentlicht: (2025)
von: Imran, Muhammad Osama, et al.
Veröffentlicht: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
Survey of Cultural Awareness in Language Models: Text and Beyond
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2024)
von: Pawar, Siddhesh, et al.
Veröffentlicht: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
von: Oh, Juhyun, et al.
Veröffentlicht: (2024)
Neuropsychology of AI: Relationship Between Activation Proximity and Categorical Proximity Within Neural Categories of Synthetic Cognition
von: Pichat, Michael, et al.
Veröffentlicht: (2024)
von: Pichat, Michael, et al.
Veröffentlicht: (2024)
DeGenTWeb: A First Look at LLM-dominant Websites
von: He, Sichang Steven, et al.
Veröffentlicht: (2026)
von: He, Sichang Steven, et al.
Veröffentlicht: (2026)
LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation
von: Park, Junyeong, et al.
Veröffentlicht: (2025)
von: Park, Junyeong, et al.
Veröffentlicht: (2025)
TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI
von: Li, Jingming
Veröffentlicht: (2025)
von: Li, Jingming
Veröffentlicht: (2025)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
von: Yan, Shuo, et al.
Veröffentlicht: (2025)
von: Yan, Shuo, et al.
Veröffentlicht: (2025)
A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist
von: Jeon, Sohyeon, et al.
Veröffentlicht: (2025)
von: Jeon, Sohyeon, et al.
Veröffentlicht: (2025)
Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
von: Jin, Jiho, et al.
Veröffentlicht: (2025)
von: Jin, Jiho, et al.
Veröffentlicht: (2025)
Improving LLM Abilities in Idiomatic Translation
von: Donthi, Sundesh, et al.
Veröffentlicht: (2024)
von: Donthi, Sundesh, et al.
Veröffentlicht: (2024)
EmoLLM: Appraisal-Grounded Cognitive-Emotional Co-Reasoning in Large Language Models
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
Grounding from an AI and Cognitive Science Lens
von: Bajaj, Goonmeet, et al.
Veröffentlicht: (2024)
von: Bajaj, Goonmeet, et al.
Veröffentlicht: (2024)
UniCog: Uncovering Cognitive Abilities of LLMs through Latent Mind Space Analysis
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation
von: Xu, Zonghuan, et al.
Veröffentlicht: (2026)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2026)
Scheming Ability in LLM-to-LLM Strategic Interactions
von: Pham, Thao
Veröffentlicht: (2025)
von: Pham, Thao
Veröffentlicht: (2025)
Evaluating the Logical Reasoning Abilities of Large Reasoning Models
von: Liu, Hanmeng, et al.
Veröffentlicht: (2025)
von: Liu, Hanmeng, et al.
Veröffentlicht: (2025)
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
von: Wan, Haiyuan, et al.
Veröffentlicht: (2025)
von: Wan, Haiyuan, et al.
Veröffentlicht: (2025)
LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation
von: Lee, Yong Oh, et al.
Veröffentlicht: (2025)
von: Lee, Yong Oh, et al.
Veröffentlicht: (2025)
Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability
von: Chen, Zhaoyu, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoyu, et al.
Veröffentlicht: (2025)
Neuropsychology and Explainability of AI: A Distributional Approach to the Relationship Between Activation Similarity of Neural Categories in Synthetic Cognition
von: Pichat, Michael, et al.
Veröffentlicht: (2024)
von: Pichat, Michael, et al.
Veröffentlicht: (2024)
Learning Compact Representations of LLM Abilities via Item Response Theory
von: Chen, Jianhao, et al.
Veröffentlicht: (2025)
von: Chen, Jianhao, et al.
Veröffentlicht: (2025)
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
von: Choukrani, Omar, et al.
Veröffentlicht: (2025)
von: Choukrani, Omar, et al.
Veröffentlicht: (2025)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
von: Huang, Jiahao, et al.
Veröffentlicht: (2026)
von: Huang, Jiahao, et al.
Veröffentlicht: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
Combating the effects of speed and delays in end-to-end self-driving
von: Tampuu, Ardi, et al.
Veröffentlicht: (2023)
von: Tampuu, Ardi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
OLA: Output Language Alignment in Code-Switched LLM Interactions
von: Oh, Juhyun, et al.
Veröffentlicht: (2026) -
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese
von: Putri, Rifki Afina, et al.
Veröffentlicht: (2024) -
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
von: Li, Zheng, et al.
Veröffentlicht: (2026) -
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024) -
Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
von: Valentin, Thomas, et al.
Veröffentlicht: (2025)