EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Anjiang, Cao, Jiannan, Li, Ran, Chen, Hongyu, Zhang, Yuhui, Wang, Ziheng, Liu, Yuan, Teixeira, Thiago S. F. X., Yang, Diyi, Wang, Ke, Aiken, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Equivalence Checking of ML GPU Kernels
by: Dubey, Kshitij, et al.
Published: (2025)
by: Dubey, Kshitij, et al.
Published: (2025)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
by: Wei, Anjiang, et al.
Published: (2024)
by: Wei, Anjiang, et al.
Published: (2024)
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
On the Semantic Expressiveness of Iso- and Equi-Recursive Types
by: Devriese, Dominique, et al.
Published: (2020)
by: Devriese, Dominique, et al.
Published: (2020)
Mixture of Small and Large Models for Chinese Spelling Check
by: Qiao, Ziheng, et al.
Published: (2025)
by: Qiao, Ziheng, et al.
Published: (2025)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
by: Jing, Huihao, et al.
Published: (2025)
by: Jing, Huihao, et al.
Published: (2025)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
by: Wang, Shengkang, et al.
Published: (2024)
by: Wang, Shengkang, et al.
Published: (2024)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
by: Wen, Bosi, et al.
Published: (2026)
by: Wen, Bosi, et al.
Published: (2026)
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
by: Wang, Qinsi, et al.
Published: (2026)
by: Wang, Qinsi, et al.
Published: (2026)
BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
DeonticBench: A Benchmark for Reasoning over Rules
by: Dou, Guangyao, et al.
Published: (2026)
by: Dou, Guangyao, et al.
Published: (2026)
CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
by: Tian, Baoliang, et al.
Published: (2025)
by: Tian, Baoliang, et al.
Published: (2025)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
Refuting Equivalence in Probabilistic Programs with Conditioning
by: Chatterjee, Krishnendu, et al.
Published: (2025)
by: Chatterjee, Krishnendu, et al.
Published: (2025)
Equivalence and Similarity Refutation for Probabilistic Programs
by: Chatterjee, Krishnendu, et al.
Published: (2024)
by: Chatterjee, Krishnendu, et al.
Published: (2024)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
by: Xiao, Ruixuan, et al.
Published: (2024)
by: Xiao, Ruixuan, et al.
Published: (2024)
Bialgebraic Reasoning on Higher-Order Program Equivalence
by: Goncharov, Sergey, et al.
Published: (2024)
by: Goncharov, Sergey, et al.
Published: (2024)
DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graph
by: Zhang, Zhehao, et al.
Published: (2024)
by: Zhang, Zhehao, et al.
Published: (2024)
ProBench: Benchmarking Large Language Models in Competitive Programming
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
Example-Based Reasoning about the Realizability of Polymorphic Programs
by: Mulleners, Niek, et al.
Published: (2024)
by: Mulleners, Niek, et al.
Published: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
L0-Reasoning Bench: Evaluating Procedural Correctness in Language Models via Simple Program Execution
by: Sun, Simeng, et al.
Published: (2025)
by: Sun, Simeng, et al.
Published: (2025)
BenchBench: Benchmarking Automated Benchmark Generation
by: Zheng, Yandan, et al.
Published: (2026)
by: Zheng, Yandan, et al.
Published: (2026)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
by: Halder, Deepon, et al.
Published: (2025)
by: Halder, Deepon, et al.
Published: (2025)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
by: Li, Yangzhuo, et al.
Published: (2026)
by: Li, Yangzhuo, et al.
Published: (2026)
Scalable Equivalence Checking and Verification of Shallow Quantum Circuits
by: Yu, Nengkun, et al.
Published: (2025)
by: Yu, Nengkun, et al.
Published: (2025)
Evaluating Program Semantics Reasoning with Type Inference in System F
by: He, Yifeng, et al.
Published: (2025)
by: He, Yifeng, et al.
Published: (2025)
Anchor Points: Benchmarking Models with Much Fewer Examples
by: Vivek, Rajan, et al.
Published: (2023)
by: Vivek, Rajan, et al.
Published: (2023)
Beyond the Phase Ordering Problem: Finding the Globally Optimal Code w.r.t. Optimization Phases
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence
by: Song, Linxin, et al.
Published: (2026)
by: Song, Linxin, et al.
Published: (2026)
SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials
by: Wan, Qixin, et al.
Published: (2025)
by: Wan, Qixin, et al.
Published: (2025)
Similar Items
-
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025) -
Equivalence Checking of ML GPU Kernels
by: Dubey, Kshitij, et al.
Published: (2025) -
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
by: Wei, Anjiang, et al.
Published: (2025) -
Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
by: Wei, Anjiang, et al.
Published: (2024) -
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025)