Watson & Holmes: A Naturalistic Benchmark for Comparing Human and LLM Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Leelawat, Thatchawin, Griffin, Lewis D |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
por: Rombaut, Benjamin, et al.
Publicado: (2024)
por: Rombaut, Benjamin, et al.
Publicado: (2024)
Temporal Context and Architecture: A Benchmark for Naturalistic EEG Decoding
por: Ergezer, Mehmet
Publicado: (2026)
por: Ergezer, Mehmet
Publicado: (2026)
Transcript of GPT-4 playing a rogue AGI in a Matrix Game
por: Griffin, Lewis D, et al.
Publicado: (2024)
por: Griffin, Lewis D, et al.
Publicado: (2024)
Adversarial Safety-Critical Scenario Generation using Naturalistic Human Driving Priors
por: Hao, Kunkun, et al.
Publicado: (2024)
por: Hao, Kunkun, et al.
Publicado: (2024)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks
por: Galatzer-Levy, Isaac R., et al.
Publicado: (2024)
por: Galatzer-Levy, Isaac R., et al.
Publicado: (2024)
ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming
por: Yang, Xinwei, et al.
Publicado: (2025)
por: Yang, Xinwei, et al.
Publicado: (2025)
Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning
por: Moreira, Benjamin Grando
Publicado: (2025)
por: Moreira, Benjamin Grando
Publicado: (2025)
ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization
por: Tso, Joseph, et al.
Publicado: (2026)
por: Tso, Joseph, et al.
Publicado: (2026)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
por: Potamitis, Nearchos, et al.
Publicado: (2025)
por: Potamitis, Nearchos, et al.
Publicado: (2025)
Benchmarking of LLM Detection: Comparing Two Competing Approaches
por: Pröhl, Thorsten, et al.
Publicado: (2024)
por: Pröhl, Thorsten, et al.
Publicado: (2024)
H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark
por: LeGris, Solim, et al.
Publicado: (2024)
por: LeGris, Solim, et al.
Publicado: (2024)
HEARTS: Benchmarking LLM Reasoning on Health Time Series
por: Li, Sirui, et al.
Publicado: (2026)
por: Li, Sirui, et al.
Publicado: (2026)
Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study
por: Goebel, Kai, et al.
Publicado: (2025)
por: Goebel, Kai, et al.
Publicado: (2025)
A Representationalist, Functionalist and Naturalistic Conception of Intelligence as a Foundation for AGI
por: Pfister, Rolf
Publicado: (2025)
por: Pfister, Rolf
Publicado: (2025)
Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
por: Parashar, Shubham, et al.
Publicado: (2025)
por: Parashar, Shubham, et al.
Publicado: (2025)
Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
por: Zhao, Tianzhe, et al.
Publicado: (2026)
por: Zhao, Tianzhe, et al.
Publicado: (2026)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
por: Li, Kuan, et al.
Publicado: (2026)
por: Li, Kuan, et al.
Publicado: (2026)
LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
por: Agarwal, Shradha, et al.
Publicado: (2026)
por: Agarwal, Shradha, et al.
Publicado: (2026)
An Analysis of Architectural Impact on LLM-based Abstract Visual Reasoning: A Systematic Benchmark on RAVEN-FAIR
por: Urgun, Sinan, et al.
Publicado: (2025)
por: Urgun, Sinan, et al.
Publicado: (2025)
A Law Reasoning Benchmark for LLM with Tree-Organized Structures including Factum Probandum, Evidence and Experiences
por: Shen, Jiaxin, et al.
Publicado: (2025)
por: Shen, Jiaxin, et al.
Publicado: (2025)
Inducing Personality in LLM-Based Honeypot Agents: Measuring the Effect on Human-Like Agenda Generation
por: Newsham, Lewis, et al.
Publicado: (2025)
por: Newsham, Lewis, et al.
Publicado: (2025)
PrivacyReasoner: Can LLM Emulate a Human-like Privacy Mind?
por: Tu, Yiwen, et al.
Publicado: (2026)
por: Tu, Yiwen, et al.
Publicado: (2026)
MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
por: Lei, Yuzhen, et al.
Publicado: (2025)
por: Lei, Yuzhen, et al.
Publicado: (2025)
SPhyR: Spatial-Physical Reasoning Benchmark on Material Distribution
por: Siedler, Philipp D.
Publicado: (2025)
por: Siedler, Philipp D.
Publicado: (2025)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
por: Kolasani, Sai, et al.
Publicado: (2025)
por: Kolasani, Sai, et al.
Publicado: (2025)
Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
por: Lu, Junjie, et al.
Publicado: (2025)
por: Lu, Junjie, et al.
Publicado: (2025)
ReasoningRec: Bridging Personalized Recommendations and Human-Interpretable Explanations through LLM Reasoning
por: Bismay, Millennium, et al.
Publicado: (2024)
por: Bismay, Millennium, et al.
Publicado: (2024)
OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
por: Liu, Wanhao, et al.
Publicado: (2026)
por: Liu, Wanhao, et al.
Publicado: (2026)
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
por: Ozeki, Kentaro, et al.
Publicado: (2025)
por: Ozeki, Kentaro, et al.
Publicado: (2025)
Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
por: Kharlamova, Arina, et al.
Publicado: (2025)
por: Kharlamova, Arina, et al.
Publicado: (2025)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
por: Lin, Wenze, et al.
Publicado: (2026)
por: Lin, Wenze, et al.
Publicado: (2026)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
por: Tran, Khanh-Tung, et al.
Publicado: (2025)
por: Tran, Khanh-Tung, et al.
Publicado: (2025)
PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
por: Zhang, Xinyu, et al.
Publicado: (2025)
por: Zhang, Xinyu, et al.
Publicado: (2025)
CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
por: Cai, Jie, et al.
Publicado: (2025)
por: Cai, Jie, et al.
Publicado: (2025)
Formalization of Dialogue in the Decision Support System of Dr. Watson Type
por: Goldberg, Saveli, et al.
Publicado: (2024)
por: Goldberg, Saveli, et al.
Publicado: (2024)
AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
Mind the Gap: Evaluating the Representativeness of Quantitative Medical Language Reasoning LLM Benchmarks for African Disease Burdens
por: Mutisya, Fred, et al.
Publicado: (2025)
por: Mutisya, Fred, et al.
Publicado: (2025)
Comparative Analysis and Parametric Tuning of PPO, GRPO, and DAPO for LLM Reasoning Enhancement
por: Lian, Yongsheng
Publicado: (2025)
por: Lian, Yongsheng
Publicado: (2025)
Ejemplares similares
-
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
por: Rombaut, Benjamin, et al.
Publicado: (2024) -
Temporal Context and Architecture: A Benchmark for Naturalistic EEG Decoding
por: Ergezer, Mehmet
Publicado: (2026) -
Transcript of GPT-4 playing a rogue AGI in a Matrix Game
por: Griffin, Lewis D, et al.
Publicado: (2024) -
Adversarial Safety-Critical Scenario Generation using Naturalistic Human Driving Priors
por: Hao, Kunkun, et al.
Publicado: (2024) -
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
por: Paglieri, Davide, et al.
Publicado: (2024)