A Compute-Matched Re-Evaluation of TroVE on MATH
Fuente:
arXiv
Saved in:
| Main Authors: | Sesterhenn, Tobias, Berlot-Attwell, Ian, Zenkner, Janis, Bartelt, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms
by: Zenkner, Janis, et al.
Published: (2025)
by: Zenkner, Janis, et al.
Published: (2025)
AbstractBeam: Enhancing Bottom-Up Program Synthesis using Library Learning
by: Zenkner, Janis, et al.
Published: (2024)
by: Zenkner, Janis, et al.
Published: (2024)
Shedding Light in Task Decomposition in Program Synthesis: The Driving Force of the Synthesizer Model
by: Zenkner, Janis, et al.
Published: (2025)
by: Zenkner, Janis, et al.
Published: (2025)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
by: Wang, Zhiruo, et al.
Published: (2024)
by: Wang, Zhiruo, et al.
Published: (2024)
MCP4IFC: IFC-Based Building Design Using Large Language Models
by: Nithyanantham, Bharathi Kannan, et al.
Published: (2025)
by: Nithyanantham, Bharathi Kannan, et al.
Published: (2025)
LLM Library Learning Fails: A LEGO-Prover Case Study
by: Berlot-Attwell, Ian, et al.
Published: (2025)
by: Berlot-Attwell, Ian, et al.
Published: (2025)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
by: Berlot-Attwell, Ian, et al.
Published: (2024)
by: Berlot-Attwell, Ian, et al.
Published: (2024)
EvolVE: Evolutionary Search for LLM-based Verilog Generation and Optimization
by: Hsin, Wei-Po, et al.
Published: (2026)
by: Hsin, Wei-Po, et al.
Published: (2026)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
by: Chernyshev, Konstantin, et al.
Published: (2024)
by: Chernyshev, Konstantin, et al.
Published: (2024)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text
by: Lin, Zhenru, et al.
Published: (2024)
by: Lin, Zhenru, et al.
Published: (2024)
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
by: Singhvi, Arnav, et al.
Published: (2023)
by: Singhvi, Arnav, et al.
Published: (2023)
Turn: A Language for Agentic Computation
by: Kizito, Muyukani
Published: (2026)
by: Kizito, Muyukani
Published: (2026)
VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
by: Zeng, Lingfei, et al.
Published: (2025)
by: Zeng, Lingfei, et al.
Published: (2025)
Phyelds: A Pythonic Framework for Aggregate Computing
by: Aguzzi, Gianluca, et al.
Published: (2026)
by: Aguzzi, Gianluca, et al.
Published: (2026)
Evaluating adaptive and generative AI-based feedback and recommendations in a knowledge-graph-integrated programming learning system
by: Nongkhai, Lalita Na, et al.
Published: (2026)
by: Nongkhai, Lalita Na, et al.
Published: (2026)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
by: Cao, Jialun, et al.
Published: (2025)
by: Cao, Jialun, et al.
Published: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
by: Huang, Kaixuan, et al.
Published: (2025)
by: Huang, Kaixuan, et al.
Published: (2025)
DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation
by: Fang, Jingzhi, et al.
Published: (2026)
by: Fang, Jingzhi, et al.
Published: (2026)
Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation
by: Ding, Yaoyao, et al.
Published: (2025)
by: Ding, Yaoyao, et al.
Published: (2025)
Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
by: Wong, Lionel, et al.
Published: (2025)
by: Wong, Lionel, et al.
Published: (2025)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
by: Stengel-Eskin, Elias, et al.
Published: (2024)
by: Stengel-Eskin, Elias, et al.
Published: (2024)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
by: Saha, Soumadeep, et al.
Published: (2025)
by: Saha, Soumadeep, et al.
Published: (2025)
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024)
by: Gulati, Aryan, et al.
Published: (2024)
A Distribution Semantics for Probabilistic Term Rewriting
by: Vidal, Germán
Published: (2024)
by: Vidal, Germán
Published: (2024)
PDL: A Declarative Prompt Programming Language
by: Vaziri, Mandana, et al.
Published: (2024)
by: Vaziri, Mandana, et al.
Published: (2024)
Disentangling Exploration of Large Language Models by Optimal Exploitation
by: Grams, Tim, et al.
Published: (2025)
by: Grams, Tim, et al.
Published: (2025)
C2RUST-BENCH: A Minimized, Representative Dataset for C-to-Rust Transpilation Evaluation
by: Sirlanci, Melih, et al.
Published: (2025)
by: Sirlanci, Melih, et al.
Published: (2025)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants
by: Bayazıt, Barış, et al.
Published: (2025)
by: Bayazıt, Barış, et al.
Published: (2025)
ZeroML: A Next Generation AutoML Language
by: Mahmud, Monirul Islam
Published: (2025)
by: Mahmud, Monirul Islam
Published: (2025)
Oracular Programming: A Modular Foundation for Building LLM-Enabled Software
by: Laurent, Jonathan, et al.
Published: (2025)
by: Laurent, Jonathan, et al.
Published: (2025)
Hey Pentti, We Did It!: A Fully Vector-Symbolic Lisp
by: Tomkins-Flanagan, Eilene, et al.
Published: (2025)
by: Tomkins-Flanagan, Eilene, et al.
Published: (2025)
CodeMind: Evaluating Large Language Models for Code Reasoning
by: Liu, Changshu, et al.
Published: (2024)
by: Liu, Changshu, et al.
Published: (2024)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
by: Agarwal, Anmol, et al.
Published: (2026)
by: Agarwal, Anmol, et al.
Published: (2026)
MTP: A Meaning-Typed Language Abstraction for AI-Integrated Programming
by: Dantanarayana, Jayanaka L., et al.
Published: (2024)
by: Dantanarayana, Jayanaka L., et al.
Published: (2024)
MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark
by: Zou, Qingyun, et al.
Published: (2026)
by: Zou, Qingyun, et al.
Published: (2026)
SLaDe: A Portable Small Language Model Decompiler for Optimized Assembly
by: Armengol-Estapé, Jordi, et al.
Published: (2023)
by: Armengol-Estapé, Jordi, et al.
Published: (2023)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025)
by: Le-Cong, Thanh, et al.
Published: (2025)
Similar Items
-
Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms
by: Zenkner, Janis, et al.
Published: (2025) -
AbstractBeam: Enhancing Bottom-Up Program Synthesis using Library Learning
by: Zenkner, Janis, et al.
Published: (2024) -
Shedding Light in Task Decomposition in Program Synthesis: The Driving Force of the Synthesizer Model
by: Zenkner, Janis, et al.
Published: (2025) -
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
by: Wang, Zhiruo, et al.
Published: (2024) -
MCP4IFC: IFC-Based Building Design Using Large Language Models
by: Nithyanantham, Bharathi Kannan, et al.
Published: (2025)