AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Wei, Ma, Shuoyoucheng, Wang, Zhenhua, Wang, Enze, Chen, Kai, Sun, Xiaobing, Wang, Baosheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AIPsychoBench: Understanding the Psychometric Differences Between LLMs and Humans
by: Wei Xie, et al.
Published: (2026)
by: Wei Xie, et al.
Published: (2026)
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
by: Xie, Wei, et al.
Published: (2024)
by: Xie, Wei, et al.
Published: (2024)
Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology
by: Wang, Zhenhua, et al.
Published: (2024)
by: Wang, Zhenhua, et al.
Published: (2024)
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
by: Guo, Qianhong, et al.
Published: (2025)
by: Guo, Qianhong, et al.
Published: (2025)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)
by: Zhang, Mengyuan, et al.
Published: (2024)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
by: Liu, Zhiqiang, et al.
Published: (2025)
by: Liu, Zhiqiang, et al.
Published: (2025)
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
by: Yang, Lin, et al.
Published: (2026)
by: Yang, Lin, et al.
Published: (2026)
Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
by: Xu, Jingao, et al.
Published: (2025)
by: Xu, Jingao, et al.
Published: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
by: Tan, Haoran, et al.
Published: (2025)
by: Tan, Haoran, et al.
Published: (2025)
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
by: Zhao, Taibiao, et al.
Published: (2025)
by: Zhao, Taibiao, et al.
Published: (2025)
HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
by: Qin, Lixiong, et al.
Published: (2025)
by: Qin, Lixiong, et al.
Published: (2025)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
by: Ou, Jiao, et al.
Published: (2023)
by: Ou, Jiao, et al.
Published: (2023)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
by: Russo, Giuseppe, et al.
Published: (2025)
by: Russo, Giuseppe, et al.
Published: (2025)
Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models
by: Ye, Haoran, et al.
Published: (2024)
by: Ye, Haoran, et al.
Published: (2024)
Understanding the Collapse of LLMs in Model Editing
by: Yang, Wanli, et al.
Published: (2024)
by: Yang, Wanli, et al.
Published: (2024)
Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models
by: He-Yueya, Joy, et al.
Published: (2024)
by: He-Yueya, Joy, et al.
Published: (2024)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
Human Psychometric Questionnaires Mischaracterize LLM Behavior
by: Song, Woojung, et al.
Published: (2025)
by: Song, Woojung, et al.
Published: (2025)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
by: Li, Youquan, et al.
Published: (2024)
by: Li, Youquan, et al.
Published: (2024)
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
by: Ruan, Kai, et al.
Published: (2024)
by: Ruan, Kai, et al.
Published: (2024)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
by: Xu, Ruiling, et al.
Published: (2025)
by: Xu, Ruiling, et al.
Published: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
Rethinking the Understanding Ability across LLMs through Mutual Information
by: Wang, Shaojie, et al.
Published: (2025)
by: Wang, Shaojie, et al.
Published: (2025)
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)
by: Ren, Yanwei, et al.
Published: (2025)
Orchestrating LLMs with Different Personalizations
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
InductionBench: LLMs Fail in the Simplest Complexity Class
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment
by: Wang, Yixuan, et al.
Published: (2026)
by: Wang, Yixuan, et al.
Published: (2026)
Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
by: Lee, Seungbeen, et al.
Published: (2024)
by: Lee, Seungbeen, et al.
Published: (2024)
Probing the "Psyche'' of Large Reasoning Models: Understanding Through a Human Lens
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023)
by: Wang, Yubo, et al.
Published: (2023)
CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models
by: Zhang, Wenjing, et al.
Published: (2024)
by: Zhang, Wenjing, et al.
Published: (2024)
Dr. SoW: Density Ratio of Strong-over-weak LLMs for Reducing the Cost of Human Annotation in Preference Tuning
by: Xu, Guangxuan, et al.
Published: (2024)
by: Xu, Guangxuan, et al.
Published: (2024)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
A Closer Look into LLMs for Table Understanding
by: Wang, Jia, et al.
Published: (2026)
by: Wang, Jia, et al.
Published: (2026)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
by: Li, Zizhen, et al.
Published: (2025)
by: Li, Zizhen, et al.
Published: (2025)
Similar Items
-
AIPsychoBench: Understanding the Psychometric Differences Between LLMs and Humans
by: Wei Xie, et al.
Published: (2026) -
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
by: Xie, Wei, et al.
Published: (2024) -
Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology
by: Wang, Zhenhua, et al.
Published: (2024) -
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
by: Guo, Qianhong, et al.
Published: (2025) -
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)