Saved in:
| Main Authors: | Xiao, Xiao, Noh, Hayoun, Gonzalez-Franco, Mar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.08549 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How AI Companionship Develops: Evidence from a Longitudinal Study
by: Hwang, Angel Hsing-Chi, et al.
Published: (2025)
by: Hwang, Angel Hsing-Chi, et al.
Published: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025)
by: Liu, Hongtao, et al.
Published: (2025)
Digital Companionship: Overlapping Uses of AI Companions and AI Assistants
by: Manoli, Aikaterina, et al.
Published: (2025)
by: Manoli, Aikaterina, et al.
Published: (2025)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
by: Conde, Javier, et al.
Published: (2025)
by: Conde, Javier, et al.
Published: (2025)
Evaluating LLMs Capabilities Towards Understanding Social Dynamics
by: Tahir, Anique, et al.
Published: (2024)
by: Tahir, Anique, et al.
Published: (2024)
An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
by: Vyas, Kaustubh, et al.
Published: (2025)
by: Vyas, Kaustubh, et al.
Published: (2025)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
by: Kwon, Deuksin, et al.
Published: (2024)
by: Kwon, Deuksin, et al.
Published: (2024)
Lilith: Developmental Modular LLMs with Chemical Signaling
by: Farooqi, Mohid, et al.
Published: (2025)
by: Farooqi, Mohid, et al.
Published: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
A Comprehensive Evaluation of Cognitive Biases in LLMs
by: Malberg, Simon, et al.
Published: (2024)
by: Malberg, Simon, et al.
Published: (2024)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
by: Zhang, Shimao, et al.
Published: (2025)
by: Zhang, Shimao, et al.
Published: (2025)
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
by: Li, Ziyu, et al.
Published: (2024)
by: Li, Ziyu, et al.
Published: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
by: Zhang, Enming, et al.
Published: (2025)
by: Zhang, Enming, et al.
Published: (2025)
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
by: Sun, Zhaoyue, et al.
Published: (2026)
by: Sun, Zhaoyue, et al.
Published: (2026)
LLMs with Personalities in Multi-issue Negotiation Games
by: Noh, Sean, et al.
Published: (2024)
by: Noh, Sean, et al.
Published: (2024)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
by: Ruan, Kai, et al.
Published: (2024)
by: Ruan, Kai, et al.
Published: (2024)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off
by: Zhang, Longyin, et al.
Published: (2026)
by: Zhang, Longyin, et al.
Published: (2026)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025)
by: Rutar, Danaja, et al.
Published: (2025)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
Evidence of Cognitive Deficits andDevelopmental Advances in Generative AI: A Clock Drawing Test Analysis
by: Galatzer-Levy, Isaac R., et al.
Published: (2024)
by: Galatzer-Levy, Isaac R., et al.
Published: (2024)
Recent Advancement of Emotion Cognition in Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation
by: Sakabe, Ritsu, et al.
Published: (2025)
by: Sakabe, Ritsu, et al.
Published: (2025)
AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
by: Shi, Jinghang, et al.
Published: (2025)
by: Shi, Jinghang, et al.
Published: (2025)
Everyday AR through AI-in-the-Loop
by: Suzuki, Ryo, et al.
Published: (2024)
by: Suzuki, Ryo, et al.
Published: (2024)
Capability Self-Assessment: Teaching LLMs to Know Their Limits
by: Yang, Haoyan, et al.
Published: (2026)
by: Yang, Haoyan, et al.
Published: (2026)
Why the Valuable Capabilities of LLMs Are Precisely the Unexplainable Ones
by: Cheng, Quan
Published: (2026)
by: Cheng, Quan
Published: (2026)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Evaluating GPT's Capability in Identifying Stages of Cognitive Impairment from Electronic Health Data
by: Leng, Yu, et al.
Published: (2025)
by: Leng, Yu, et al.
Published: (2025)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
by: Li, Haoming, et al.
Published: (2025)
by: Li, Haoming, et al.
Published: (2025)
The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks
by: Galatzer-Levy, Isaac R., et al.
Published: (2024)
by: Galatzer-Levy, Isaac R., et al.
Published: (2024)
Similar Items
-
How AI Companionship Develops: Evidence from a Longitudinal Study
by: Hwang, Angel Hsing-Chi, et al.
Published: (2025) -
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025) -
Digital Companionship: Overlapping Uses of AI Companions and AI Assistants
by: Manoli, Aikaterina, et al.
Published: (2025) -
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024) -
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
by: Conde, Javier, et al.
Published: (2025)