Are LLMs Smarter Than Chimpanzees? An Evaluation on Perspective Taking and Knowledge State Estimation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Dingyi, Zhao, Junqi, Li, Xue, Li, Ce, Li, Boyang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Are Large Language Models Truly Smarter Than Humans?
par: M, Eshwar Reddy, et autres
Publié: (2026)
par: M, Eshwar Reddy, et autres
Publié: (2026)
On the Difficulty of Learning a Meta-network for Training Data Selection
par: Du, Zilin, et autres
Publié: (2026)
par: Du, Zilin, et autres
Publié: (2026)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
par: Li, Ang, et autres
Publié: (2025)
par: Li, Ang, et autres
Publié: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
par: Ballon, Marthe, et autres
Publié: (2026)
par: Ballon, Marthe, et autres
Publié: (2026)
Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
par: Yang, Zhuoyi, et autres
Publié: (2025)
par: Yang, Zhuoyi, et autres
Publié: (2025)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
par: Karan, Aayush, et autres
Publié: (2025)
par: Karan, Aayush, et autres
Publié: (2025)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
par: Kang, Yipeng, et autres
Publié: (2024)
par: Kang, Yipeng, et autres
Publié: (2024)
Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees
par: Ma, Xiaoyu, et autres
Publié: (2026)
par: Ma, Xiaoyu, et autres
Publié: (2026)
MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
par: Chen, Jiawei, et autres
Publié: (2025)
par: Chen, Jiawei, et autres
Publié: (2025)
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation
par: Zhang, Wenyu, et autres
Publié: (2024)
par: Zhang, Wenyu, et autres
Publié: (2024)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
par: Han, Chao, et autres
Publié: (2025)
par: Han, Chao, et autres
Publié: (2025)
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
par: Thiyagarajan, Prameshwar, et autres
Publié: (2025)
par: Thiyagarajan, Prameshwar, et autres
Publié: (2025)
Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
par: Zhang, Bingjie, et autres
Publié: (2025)
par: Zhang, Bingjie, et autres
Publié: (2025)
Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation
par: Li, Tianjiao, et autres
Publié: (2026)
par: Li, Tianjiao, et autres
Publié: (2026)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
par: Li, Yanhong, et autres
Publié: (2025)
par: Li, Yanhong, et autres
Publié: (2025)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
par: Liang, Yiming, et autres
Publié: (2025)
par: Liang, Yiming, et autres
Publié: (2025)
KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing
par: Li, Zhifei, et autres
Publié: (2025)
par: Li, Zhifei, et autres
Publié: (2025)
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
par: Chen, Yuanteng, et autres
Publié: (2025)
par: Chen, Yuanteng, et autres
Publié: (2025)
Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge Integration
par: Sun, Wenju, et autres
Publié: (2025)
par: Sun, Wenju, et autres
Publié: (2025)
Node Importance Estimation Leveraging LLMs for Semantic Augmentation in Knowledge Graphs
par: Lin, Xinyu, et autres
Publié: (2024)
par: Lin, Xinyu, et autres
Publié: (2024)
LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning
par: Hua, Rui, et autres
Publié: (2026)
par: Hua, Rui, et autres
Publié: (2026)
Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge Conflicts
par: Sun, Wenju, et autres
Publié: (2025)
par: Sun, Wenju, et autres
Publié: (2025)
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
par: Luo, Ruilin, et autres
Publié: (2024)
par: Luo, Ruilin, et autres
Publié: (2024)
Can LLMs Support Medical Knowledge Imputation? An Evaluation-Based Perspective
par: Yao, Xinyu, et autres
Publié: (2025)
par: Yao, Xinyu, et autres
Publié: (2025)
A Training Data Recipe to Accelerate A* Search with Language Models
par: Gupta, Devaansh, et autres
Publié: (2024)
par: Gupta, Devaansh, et autres
Publié: (2024)
BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference
par: Zhao, Junqi, et autres
Publié: (2024)
par: Zhao, Junqi, et autres
Publié: (2024)
H2HTalk: Evaluating Large Language Models as Emotional Companion
par: Wang, Boyang, et autres
Publié: (2025)
par: Wang, Boyang, et autres
Publié: (2025)
Graph Neural Networks Are More Than Filters: Revisiting and Benchmarking from A Spectral Perspective
par: Dong, Yushun, et autres
Publié: (2024)
par: Dong, Yushun, et autres
Publié: (2024)
Look Within, Why LLMs Hallucinate: A Causal Perspective
par: Li, He, et autres
Publié: (2024)
par: Li, He, et autres
Publié: (2024)
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
par: Wang, Yang, et autres
Publié: (2025)
par: Wang, Yang, et autres
Publié: (2025)
Who Do LLMs Trust? Human Experts Matter More Than Other LLMs
par: Bajaj, Anooshka, et autres
Publié: (2026)
par: Bajaj, Anooshka, et autres
Publié: (2026)
Exploring Adversarial Robustness of Deep State Space Models
par: Qi, Biqing, et autres
Publié: (2024)
par: Qi, Biqing, et autres
Publié: (2024)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
par: Pala, Tej Deep, et autres
Publié: (2025)
par: Pala, Tej Deep, et autres
Publié: (2025)
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
par: Góral, Gracjan, et autres
Publié: (2025)
par: Góral, Gracjan, et autres
Publié: (2025)
Automated Clinical Data Extraction with Knowledge Conditioned LLMs
par: Li, Diya, et autres
Publié: (2024)
par: Li, Diya, et autres
Publié: (2024)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
par: Yu, Zishun, et autres
Publié: (2025)
par: Yu, Zishun, et autres
Publié: (2025)
Diffusion Models for Smarter UAVs: Decision-Making and Modeling
par: Emami, Yousef, et autres
Publié: (2025)
par: Emami, Yousef, et autres
Publié: (2025)
Conditioning Matters: Training Diffusion Policies is Faster Than You Think
par: Dong, Zibin, et autres
Publié: (2025)
par: Dong, Zibin, et autres
Publié: (2025)
ChimpVLM: Ethogram-Enhanced Chimpanzee Behaviour Recognition
par: Brookes, Otto, et autres
Publié: (2024)
par: Brookes, Otto, et autres
Publié: (2024)
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
par: Chen, Meifang, et autres
Publié: (2026)
par: Chen, Meifang, et autres
Publié: (2026)
Documents similaires
-
Are Large Language Models Truly Smarter Than Humans?
par: M, Eshwar Reddy, et autres
Publié: (2026) -
On the Difficulty of Learning a Meta-network for Training Data Selection
par: Du, Zilin, et autres
Publié: (2026) -
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
par: Li, Ang, et autres
Publié: (2025) -
Benchmarks Saturate When The Model Gets Smarter Than The Judge
par: Ballon, Marthe, et autres
Publié: (2026) -
Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
par: Yang, Zhuoyi, et autres
Publié: (2025)