LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rabern, Brian, Mondorf, Philipp, Plank, Barbara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
Tracing Uncertainty in Language Model "Reasoning"
von: Grünefeld, Nils, et al.
Veröffentlicht: (2026)
von: Grünefeld, Nils, et al.
Veröffentlicht: (2026)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
von: Orth, Jasmin, et al.
Veröffentlicht: (2025)
von: Orth, Jasmin, et al.
Veröffentlicht: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Do Large Language Models Excel in Complex Logical Reasoning with Formal Language?
von: Jiang, Jin, et al.
Veröffentlicht: (2025)
von: Jiang, Jin, et al.
Veröffentlicht: (2025)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
von: Cheng, Xinyuan, et al.
Veröffentlicht: (2026)
von: Cheng, Xinyuan, et al.
Veröffentlicht: (2026)
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2025)
von: Ozeki, Kentaro, et al.
Veröffentlicht: (2025)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
Logical Reasoning in Large Language Models: A Survey
von: Liu, Hanmeng, et al.
Veröffentlicht: (2025)
von: Liu, Hanmeng, et al.
Veröffentlicht: (2025)
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding
von: Li, Yanda, et al.
Veröffentlicht: (2024)
von: Li, Yanda, et al.
Veröffentlicht: (2024)
LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models
von: He, Kang, et al.
Veröffentlicht: (2025)
von: He, Kang, et al.
Veröffentlicht: (2025)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games
von: Malik, Saumya
Veröffentlicht: (2024)
von: Malik, Saumya
Veröffentlicht: (2024)
GLoRE: Evaluating Logical Reasoning of Large Language Models
von: liu, Hanmeng, et al.
Veröffentlicht: (2023)
von: liu, Hanmeng, et al.
Veröffentlicht: (2023)
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
von: Liang, Qiliang, et al.
Veröffentlicht: (2026)
von: Liang, Qiliang, et al.
Veröffentlicht: (2026)
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
von: Bruun, Sofie Helene, et al.
Veröffentlicht: (2025)
von: Bruun, Sofie Helene, et al.
Veröffentlicht: (2025)
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
von: Xu, Fangzhi, et al.
Veröffentlicht: (2023)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2023)
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
von: Hong, Ruixin, et al.
Veröffentlicht: (2023)
von: Hong, Ruixin, et al.
Veröffentlicht: (2023)
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
von: Chen, Jiangjie, et al.
Veröffentlicht: (2025)
von: Chen, Jiangjie, et al.
Veröffentlicht: (2025)
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
von: Kai, Ding, et al.
Veröffentlicht: (2024)
von: Kai, Ding, et al.
Veröffentlicht: (2024)
Structured Chemistry Reasoning with Large Language Models
von: Ouyang, Siru, et al.
Veröffentlicht: (2023)
von: Ouyang, Siru, et al.
Veröffentlicht: (2023)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
von: Luo, Man, et al.
Veröffentlicht: (2023)
von: Luo, Man, et al.
Veröffentlicht: (2023)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
von: Zhu, Yakun, et al.
Veröffentlicht: (2025)
von: Zhu, Yakun, et al.
Veröffentlicht: (2025)
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
von: Junias, Obed, et al.
Veröffentlicht: (2026)
von: Junias, Obed, et al.
Veröffentlicht: (2026)
ChatRule: Mining Logical Rules with Large Language Models for Knowledge Graph Reasoning
von: Luo, Linhao, et al.
Veröffentlicht: (2023)
von: Luo, Linhao, et al.
Veröffentlicht: (2023)
Enhancing Large Language Models through Structured Reasoning
von: Dong, Yubo, et al.
Veröffentlicht: (2025)
von: Dong, Yubo, et al.
Veröffentlicht: (2025)
SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
von: Yang, Wanqi, et al.
Veröffentlicht: (2025)
von: Yang, Wanqi, et al.
Veröffentlicht: (2025)
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Conjecturing: An Overlooked Step in Formal Mathematical Reasoning
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2025)
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2025)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024) -
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024) -
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024) -
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025) -
Tracing Uncertainty in Language Model "Reasoning"
von: Grünefeld, Nils, et al.
Veröffentlicht: (2026)