Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason's Selection Task
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abe, Hirohiko, Ozeki, Kentaro, Ando, Risako, Morishita, Takanobu, Mineshima, Koji, Okada, Mitsuhiro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Abductive Reasoning with Syllogistic Forms in Large Language Models
von: Abe, Hirohiko, et al.
Veröffentlicht: (2026)
von: Abe, Hirohiko, et al.
Veröffentlicht: (2026)
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2025)
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2025)
Exploring Reasoning Biases in Large Language Models Through Syllogism: Insights from the NeuBAROCO Dataset
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2024)
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2024)
A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
von: Funakura, Hayate, et al.
Veröffentlicht: (2025)
von: Funakura, Hayate, et al.
Veröffentlicht: (2025)
Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
von: Ueda, Kentaro, et al.
Veröffentlicht: (2025)
von: Ueda, Kentaro, et al.
Veröffentlicht: (2025)
DeonticBench: A Benchmark for Reasoning over Rules
von: Dou, Guangyao, et al.
Veröffentlicht: (2026)
von: Dou, Guangyao, et al.
Veröffentlicht: (2026)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
An Investigation of Warning Erroneous Chat Translations in Cross-lingual Communication
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
Role-Conditioned Refusals: Evaluating Access Control Reasoning in Large Language Models
von: Klisura, Đorđe, et al.
Veröffentlicht: (2025)
von: Klisura, Đorđe, et al.
Veröffentlicht: (2025)
Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
PheMT: A Phenomenon-wise Dataset for Machine Translation Robustness on User-Generated Contents
von: Fujii, Ryo, et al.
Veröffentlicht: (2020)
von: Fujii, Ryo, et al.
Veröffentlicht: (2020)
Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
Enhancing Large Language Models with Neurosymbolic Reasoning for Multilingual Tasks
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2025)
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2025)
TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
Evaluating Large Language Models for Material Selection
von: Grandi, Daniele, et al.
Veröffentlicht: (2024)
von: Grandi, Daniele, et al.
Veröffentlicht: (2024)
Evaluating Multi-Hop Reasoning in Large Language Models: A Chemistry-Centric Case Study
von: Khodadad, Mohammad, et al.
Veröffentlicht: (2025)
von: Khodadad, Mohammad, et al.
Veröffentlicht: (2025)
From Monolingual to Bilingual: Investigating Language Conditioning in Large Language Models for Psycholinguistic Tasks
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yuan, Shuzhou, et al.
Veröffentlicht: (2025)
Aligning Large Language Model Behavior with Human Citation Preferences
von: Ando, Kenichiro, et al.
Veröffentlicht: (2026)
von: Ando, Kenichiro, et al.
Veröffentlicht: (2026)
Reasoning Capabilities of Large Language Models on Dynamic Tasks
von: Wong, Annie, et al.
Veröffentlicht: (2025)
von: Wong, Annie, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Math Reasoning Tasks
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
Evaluating Counterfactual Strategic Reasoning in Large Language Models
von: Georgousis, Dimitrios, et al.
Veröffentlicht: (2026)
von: Georgousis, Dimitrios, et al.
Veröffentlicht: (2026)
Evaluating Accounting Reasoning Capabilities of Large Language Models
von: Zhou, Jie, et al.
Veröffentlicht: (2026)
von: Zhou, Jie, et al.
Veröffentlicht: (2026)
GroundCocoa: A Benchmark for Evaluating Compositional & Conditional Reasoning in Language Models
von: Kohli, Harsh, et al.
Veröffentlicht: (2024)
von: Kohli, Harsh, et al.
Veröffentlicht: (2024)
Reliable Control-Point Selection for Steering Reasoning in Large Language Models
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
von: Wang, Yajing, et al.
Veröffentlicht: (2024)
von: Wang, Yajing, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in Theory of Mind Tasks
von: Kosinski, Michal
Veröffentlicht: (2023)
von: Kosinski, Michal
Veröffentlicht: (2023)
A Systematic Evaluation of Large Language Models for Natural Language Generation Tasks
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation
von: Tonja, Atnafu Lambebo, et al.
Veröffentlicht: (2024)
von: Tonja, Atnafu Lambebo, et al.
Veröffentlicht: (2024)
Evaluating Computational Accuracy of Large Language Models in Numerical Reasoning Tasks for Healthcare Applications
von: Malghan, Arjun R.
Veröffentlicht: (2025)
von: Malghan, Arjun R.
Veröffentlicht: (2025)
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
von: Hirano, Masanori, et al.
Veröffentlicht: (2024)
von: Hirano, Masanori, et al.
Veröffentlicht: (2024)
Refactoring Programs Using Large Language Models with Few-Shot Examples
von: Shirafuji, Atsushi, et al.
Veröffentlicht: (2023)
von: Shirafuji, Atsushi, et al.
Veröffentlicht: (2023)
EconNLI: Evaluating Large Language Models on Economics Reasoning
von: Guo, Yue, et al.
Veröffentlicht: (2024)
von: Guo, Yue, et al.
Veröffentlicht: (2024)
The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models
von: Lee, Taewhoo, et al.
Veröffentlicht: (2025)
von: Lee, Taewhoo, et al.
Veröffentlicht: (2025)
Evaluating Ill-Defined Tasks in Large Language Models
von: Zhou, Yi, et al.
Veröffentlicht: (2026)
von: Zhou, Yi, et al.
Veröffentlicht: (2026)
DrawingBench: Evaluating Spatial Reasoning and UI Interaction Capabilities of Large Language Models through Mouse-Based Drawing Tasks
von: Kim, Hyunjun, et al.
Veröffentlicht: (2025)
von: Kim, Hyunjun, et al.
Veröffentlicht: (2025)
Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models
von: Figueras, Banca Calvo, et al.
Veröffentlicht: (2025)
von: Figueras, Banca Calvo, et al.
Veröffentlicht: (2025)
Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese
von: Kawai, Masataka, et al.
Veröffentlicht: (2026)
von: Kawai, Masataka, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Abductive Reasoning with Syllogistic Forms in Large Language Models
von: Abe, Hirohiko, et al.
Veröffentlicht: (2026) -
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2025) -
Exploring Reasoning Biases in Large Language Models Through Syllogism: Insights from the NeuBAROCO Dataset
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2024) -
A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
von: Funakura, Hayate, et al.
Veröffentlicht: (2025) -
Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
von: Ueda, Kentaro, et al.
Veröffentlicht: (2025)