SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Anjiang, Wu, Yuheng, Wan, Yingjia, Suresh, Tarun, Tan, Huanmi, Zhou, Zhanke, Koyejo, Sanmi, Wang, Ke, Aiken, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Logic Optimization Meets SAT: A Novel Framework for Circuit-SAT Solving
by: Shi, Zhengyuan, et al.
Published: (2024)
by: Shi, Zhengyuan, et al.
Published: (2024)
SAT-based Learning of Computation Tree Logic
by: Pommellet, Adrien, et al.
Published: (2024)
by: Pommellet, Adrien, et al.
Published: (2024)
Evaluating SAT and SMT Solvers on Large-Scale Sudoku Puzzles
by: Davis, Liam, et al.
Published: (2025)
by: Davis, Liam, et al.
Published: (2025)
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Synthesis Benchmarks for Automated Reasoning
by: Hajdu, Márton, et al.
Published: (2025)
by: Hajdu, Márton, et al.
Published: (2025)
Can Transformers Reason Logically? A Study in SAT Solving
by: Pan, Leyan, et al.
Published: (2024)
by: Pan, Leyan, et al.
Published: (2024)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Pantograph: A Machine-to-Machine Interaction Interface for Advanced Theorem Proving, High Level Reasoning, and Data Extraction in Lean 4
by: Aniva, Leni, et al.
Published: (2024)
by: Aniva, Leni, et al.
Published: (2024)
The Complexity of Defining and Separating Fixpoint Formulae in Modal Logic
by: Jung, Jean Christoph, et al.
Published: (2025)
by: Jung, Jean Christoph, et al.
Published: (2025)
Modal Logic for Reasoning About Uncertainty and Confusion
by: Bílková, Marta, et al.
Published: (2025)
by: Bílková, Marta, et al.
Published: (2025)
Dsat: A Native SAT Solver for Discrete Logic
by: Zhang, Yaofang, et al.
Published: (2026)
by: Zhang, Yaofang, et al.
Published: (2026)
Identifying and Explaining (Non-)Equivalence of First-Order Logic Formulas
by: Vehlken, Fabian, et al.
Published: (2026)
by: Vehlken, Fabian, et al.
Published: (2026)
Tableaux for Automated Reasoning in Dependently-Typed Higher-Order Logic (Extended Version)
by: Niederhauser, Johannes, et al.
Published: (2024)
by: Niederhauser, Johannes, et al.
Published: (2024)
Deeply Optimizing the SAT Solver for the IC3 Algorithm
by: Su, Yuheng, et al.
Published: (2025)
by: Su, Yuheng, et al.
Published: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Equivalence and Conditional Independence in Atomic Sheaf Logic
by: Simpson, Alex
Published: (2024)
by: Simpson, Alex
Published: (2024)
PolySAT: Word-level Bit-vector Reasoning in Z3
by: Rath, Jakob, et al.
Published: (2024)
by: Rath, Jakob, et al.
Published: (2024)
RustSAT: A Library For SAT Solving in Rust
by: Jabs, Christoph
Published: (2025)
by: Jabs, Christoph
Published: (2025)
A Game for Counting Logic Formula Size and an Application to Linear Orders
by: Fournier, Gregoire, et al.
Published: (2025)
by: Fournier, Gregoire, et al.
Published: (2025)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
by: Rahman, A M Muntasir, et al.
Published: (2024)
by: Rahman, A M Muntasir, et al.
Published: (2024)
Semi-Substructural Logics à la Lambek
by: Wan, Cheng-Syuan
Published: (2024)
by: Wan, Cheng-Syuan
Published: (2024)
Reducing Arbitrary Metric Temporal Formulas into Logic Programs under Answer Set Semantics
by: Diéguez, Martín, et al.
Published: (2026)
by: Diéguez, Martín, et al.
Published: (2026)
Automatic Verification of Floating-Point Accumulation Networks
by: Zhang, David K., et al.
Published: (2025)
by: Zhang, David K., et al.
Published: (2025)
Semi-Substructural Logics with Additives
by: Veltri, Niccolò, et al.
Published: (2024)
by: Veltri, Niccolò, et al.
Published: (2024)
SAT-Based Subsumption Resolution
by: Coutelier, Robin, et al.
Published: (2024)
by: Coutelier, Robin, et al.
Published: (2024)
Life span of SAT techniques
by: Fleury, Mathias, et al.
Published: (2024)
by: Fleury, Mathias, et al.
Published: (2024)
LLM-ARC: Enhancing LLMs with an Automated Reasoning Critic
by: Kalyanpur, Aditya, et al.
Published: (2024)
by: Kalyanpur, Aditya, et al.
Published: (2024)
Agent-Knowledge Logic for Alternative Epistemic Logic
by: Nishimura, Yuki
Published: (2024)
by: Nishimura, Yuki
Published: (2024)
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
by: Gulati, Aryan, et al.
Published: (2025)
by: Gulati, Aryan, et al.
Published: (2025)
Total Outcome Logic: Unified Reasoning for a Taxonomy of Program Logics
by: Li, James, et al.
Published: (2024)
by: Li, James, et al.
Published: (2024)
Formalizing Kantian Ethics: Formula of the Universal Law Logic (FULL)
by: Olson, Taylor
Published: (2026)
by: Olson, Taylor
Published: (2026)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)
by: Zhou, Yujun, et al.
Published: (2025)
Nested Sequents for Intermediate Logics: The Case of Gödel-Dummett Logics
by: Lyon, Tim S.
Published: (2023)
by: Lyon, Tim S.
Published: (2023)
Justification Logic for Intuitionistic Modal Logic (Extended Technical Report)
by: Marin, Sonia, et al.
Published: (2025)
by: Marin, Sonia, et al.
Published: (2025)
SAT-Inspired Higher-Order Eliminations
by: Blanchette, Jasmin, et al.
Published: (2022)
by: Blanchette, Jasmin, et al.
Published: (2022)
Between proof construction and SAT-solving
by: Schubert, Aleksy, et al.
Published: (2024)
by: Schubert, Aleksy, et al.
Published: (2024)
Model Counting for Dependency Quantified Boolean Formulas
by: Fung, Long-Hin, et al.
Published: (2025)
by: Fung, Long-Hin, et al.
Published: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
Proof Complexity of Linear Logics
by: Tabatabai, Amirhossein Akbar, et al.
Published: (2026)
by: Tabatabai, Amirhossein Akbar, et al.
Published: (2026)
Similar Items
-
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025) -
Logic Optimization Meets SAT: A Novel Framework for Circuit-SAT Solving
by: Shi, Zhengyuan, et al.
Published: (2024) -
SAT-based Learning of Computation Tree Logic
by: Pommellet, Adrien, et al.
Published: (2024) -
Evaluating SAT and SMT Solvers on Large-Scale Sudoku Puzzles
by: Davis, Liam, et al.
Published: (2025) -
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
by: Wei, Anjiang, et al.
Published: (2025)