CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Anjiang, Suresh, Tarun, Cao, Jiannan, Kannan, Naveen, Wu, Yuheng, Yan, Kai, Teixeira, Thiago S. F. X., Wang, Ke, Aiken, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
by: Wei, Anjiang, et al.
Published: (2024)
by: Wei, Anjiang, et al.
Published: (2024)
Mapple: A Domain-Specific Language for Mapping Distributed Programs
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
SynCode: LLM Generation with Grammar Augmentation
by: Ugare, Shubham, et al.
Published: (2024)
by: Ugare, Shubham, et al.
Published: (2024)
Code-Driven Inductive Synthesis: Enhancing Reasoning Abilities of Large Language Models with Sequences
by: Chen, Kedi, et al.
Published: (2025)
by: Chen, Kedi, et al.
Published: (2025)
Equivalence Checking of ML GPU Kernels
by: Dubey, Kshitij, et al.
Published: (2025)
by: Dubey, Kshitij, et al.
Published: (2025)
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
by: Akgül, Ömer Faruk, et al.
Published: (2026)
by: Akgül, Ömer Faruk, et al.
Published: (2026)
CRANE: Reasoning with constrained LLM generation
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Agentic Separation Logic Specification Synthesis
by: Suresh, Tarun, et al.
Published: (2026)
by: Suresh, Tarun, et al.
Published: (2026)
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
by: 5 Team, et al.
Published: (2025)
by: 5 Team, et al.
Published: (2025)
Program Synthesis using Inductive Logic Programming for the Abstraction and Reasoning Corpus
by: Rocha, Filipe Marinho, et al.
Published: (2024)
by: Rocha, Filipe Marinho, et al.
Published: (2024)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
LLM-ARC: Enhancing LLMs with an Automated Reasoning Critic
by: Kalyanpur, Aditya, et al.
Published: (2024)
by: Kalyanpur, Aditya, et al.
Published: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
Inductive Synthesis of Inductive Heap Predicates -- Extended Version
by: Yang, Ziyi, et al.
Published: (2025)
by: Yang, Ziyi, et al.
Published: (2025)
On LLM-Based Scientific Inductive Reasoning Beyond Equations
by: Lin, Brian S., et al.
Published: (2025)
by: Lin, Brian S., et al.
Published: (2025)
LLM as Prompter: Low-resource Inductive Reasoning on Arbitrary Knowledge Graphs
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
by: Qiu, Linlu, et al.
Published: (2023)
by: Qiu, Linlu, et al.
Published: (2023)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
by: Li, Zhuohao, et al.
Published: (2025)
by: Li, Zhuohao, et al.
Published: (2025)
BEAVER: An Efficient Deterministic LLM Verifier
by: Suresh, Tarun, et al.
Published: (2025)
by: Suresh, Tarun, et al.
Published: (2025)
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
ARC-TGI: Human-Validated Task Generators with Reasoning Chain Templates for ARC-AGI
by: Lehmann, Jens, et al.
Published: (2026)
by: Lehmann, Jens, et al.
Published: (2026)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Language Models as Inductive Reasoners
by: Yang, Zonglin, et al.
Published: (2022)
by: Yang, Zonglin, et al.
Published: (2022)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Inductive Bias Extraction and Matching for LLM Prompts
by: Angel, Christian M., et al.
Published: (2025)
by: Angel, Christian M., et al.
Published: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
CausalARC: Abstract Reasoning with Causal World Models
by: Maasch, Jacqueline, et al.
Published: (2025)
by: Maasch, Jacqueline, et al.
Published: (2025)
IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking
by: Ugare, Shubham, et al.
Published: (2024)
by: Ugare, Shubham, et al.
Published: (2024)
LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning
by: Zhu, Yu, et al.
Published: (2026)
by: Zhu, Yu, et al.
Published: (2026)
ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis
by: Naik, Atharva, et al.
Published: (2026)
by: Naik, Atharva, et al.
Published: (2026)
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
by: Saha, Anisha, et al.
Published: (2025)
by: Saha, Anisha, et al.
Published: (2025)
Can Input Attributions Explain Inductive Reasoning in In-Context Learning?
by: Ye, Mengyu, et al.
Published: (2024)
by: Ye, Mengyu, et al.
Published: (2024)
Similar Items
-
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
by: Wei, Anjiang, et al.
Published: (2025) -
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
by: Wei, Anjiang, et al.
Published: (2025) -
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
by: Wei, Anjiang, et al.
Published: (2025) -
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025) -
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
by: Wei, Anjiang, et al.
Published: (2025)