A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | Funakura, Hayate, Kim, Hyunsoo, Mineshima, Koji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RLMEval: Evaluating Research-Level Neural Theorem Proving
by: Poiroux, Auguste, et al.
Published: (2025)
by: Poiroux, Auguste, et al.
Published: (2025)
Benchmarking Testing in Automated Theorem Proving
by: Kim, Jongyoon, et al.
Published: (2026)
by: Kim, Jongyoon, et al.
Published: (2026)
Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason's Selection Task
by: Abe, Hirohiko, et al.
Published: (2026)
by: Abe, Hirohiko, et al.
Published: (2026)
miniCTX: Neural Theorem Proving with (Long-)Contexts
by: Hu, Jiewen, et al.
Published: (2024)
by: Hu, Jiewen, et al.
Published: (2024)
AutoTemplate: A Simple Recipe for Lexically Constrained Text Generation
by: Iso, Hayate
Published: (2022)
by: Iso, Hayate
Published: (2022)
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
by: Ozeki, Kentaro, et al.
Published: (2025)
by: Ozeki, Kentaro, et al.
Published: (2025)
Neural Semantic Parsing with Extremely Rich Symbolic Meaning Representations
by: Zhang, Xiao, et al.
Published: (2024)
by: Zhang, Xiao, et al.
Published: (2024)
Abductive Reasoning with Syllogistic Forms in Large Language Models
by: Abe, Hirohiko, et al.
Published: (2026)
by: Abe, Hirohiko, et al.
Published: (2026)
Exploring Reasoning Biases in Large Language Models Through Syllogism: Insights from the NeuBAROCO Dataset
by: Ozeki, Kentaro, et al.
Published: (2024)
by: Ozeki, Kentaro, et al.
Published: (2024)
Automated Theorem Proving for Prolog Verification
by: Mesnard, Fred, et al.
Published: (2026)
by: Mesnard, Fred, et al.
Published: (2026)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
Aristotle: IMO-level Automated Theorem Proving
by: Achim, Tudor, et al.
Published: (2025)
by: Achim, Tudor, et al.
Published: (2025)
Scaling Natural-Language Graph-Based Test Time Compute for Automated Theorem Proving
by: Li, Vincent, et al.
Published: (2025)
by: Li, Vincent, et al.
Published: (2025)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
SCORE: A Semantic Evaluation Framework for Generative Document Parsing
by: Li, Renyu, et al.
Published: (2025)
by: Li, Renyu, et al.
Published: (2025)
BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing
by: Roy, Subhro, et al.
Published: (2022)
by: Roy, Subhro, et al.
Published: (2022)
Semantic Parsing with Candidate Expressions for Knowledge Base Question Answering
by: Nam, Daehwan, et al.
Published: (2024)
by: Nam, Daehwan, et al.
Published: (2024)
OProver: A Unified Framework for Agentic Formal Theorem Proving
by: Ma, David, et al.
Published: (2026)
by: Ma, David, et al.
Published: (2026)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
PhysProver: Advancing Automatic Theorem Proving for Physics
by: Zhang, Hanning, et al.
Published: (2026)
by: Zhang, Hanning, et al.
Published: (2026)
PMB5: Gaining More Insight into Neural Semantic Parsing with Challenging Benchmarks
by: Zhang, Xiao, et al.
Published: (2024)
by: Zhang, Xiao, et al.
Published: (2024)
Mechanic: Sorrifier-Driven Formal Decomposition Workflow for Automated Theorem Proving
by: Qiu, Ruichen, et al.
Published: (2026)
by: Qiu, Ruichen, et al.
Published: (2026)
Learning to Reason with Insight for Informal Theorem Proving
by: Li, Yunhe, et al.
Published: (2026)
by: Li, Yunhe, et al.
Published: (2026)
Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing
by: Kang, Deokhyung, et al.
Published: (2024)
by: Kang, Deokhyung, et al.
Published: (2024)
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
by: Liu, Chengwu, et al.
Published: (2026)
by: Liu, Chengwu, et al.
Published: (2026)
Neural Theorem Proving: Generating and Structuring Proofs for Formal Verification
by: Rao, Balaji, et al.
Published: (2025)
by: Rao, Balaji, et al.
Published: (2025)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
by: Quan, Xin, et al.
Published: (2025)
by: Quan, Xin, et al.
Published: (2025)
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
by: Chen, Luoxin, et al.
Published: (2025)
by: Chen, Luoxin, et al.
Published: (2025)
Mathematical Formalized Problem Solving and Theorem Proving in Different Fields in Lean 4
by: Tang, Xichen
Published: (2024)
by: Tang, Xichen
Published: (2024)
Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving
by: Quan, Xin, et al.
Published: (2024)
by: Quan, Xin, et al.
Published: (2024)
Cleanse: Uncertainty Estimation Approach Using Clustering-based Semantic Consistency in LLMs
by: Joo, Minsuh, et al.
Published: (2025)
by: Joo, Minsuh, et al.
Published: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
by: Zhang, Ziyin, et al.
Published: (2025)
by: Zhang, Ziyin, et al.
Published: (2025)
Exploring In-Context Learning for Frame-Semantic Parsing
by: Garat, Diego, et al.
Published: (2025)
by: Garat, Diego, et al.
Published: (2025)
Scope-enhanced Compositional Semantic Parsing for DRT
by: Yang, Xiulin, et al.
Published: (2024)
by: Yang, Xiulin, et al.
Published: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024)
by: Niwa, Ayana, et al.
Published: (2024)
Compiling by Proving: Language-Agnostic Automatic Optimization from Formal Semantics
by: Zhao, Jianhong, et al.
Published: (2025)
by: Zhao, Jianhong, et al.
Published: (2025)
MerLean-Prover: A Recursive Looping Harness for Lean 4 Theorem Proving
by: Li, Jinzheng, et al.
Published: (2026)
by: Li, Jinzheng, et al.
Published: (2026)
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
by: Iso, Hayate, et al.
Published: (2025)
by: Iso, Hayate, et al.
Published: (2025)
Spark-Prover-X1: Formal Theorem Proving Through Diverse Data Training
by: Zhou, Xinyuan, et al.
Published: (2025)
by: Zhou, Xinyuan, et al.
Published: (2025)
Similar Items
-
RLMEval: Evaluating Research-Level Neural Theorem Proving
by: Poiroux, Auguste, et al.
Published: (2025) -
Benchmarking Testing in Automated Theorem Proving
by: Kim, Jongyoon, et al.
Published: (2026) -
Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason's Selection Task
by: Abe, Hirohiko, et al.
Published: (2026) -
miniCTX: Neural Theorem Proving with (Long-)Contexts
by: Hu, Jiewen, et al.
Published: (2024) -
AutoTemplate: A Simple Recipe for Lexically Constrained Text Generation
by: Iso, Hayate
Published: (2022)