A Compute-Matched Re-Evaluation of TroVE on MATH

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sesterhenn, Tobias, Berlot-Attwell, Ian, Zenkner, Janis, Bartelt, Christian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915418040434688
author Sesterhenn, Tobias
Berlot-Attwell, Ian
Zenkner, Janis
Bartelt, Christian
author_facet Sesterhenn, Tobias
Berlot-Attwell, Ian
Zenkner, Janis
Bartelt, Christian
contents Reusing established theorems and formulas is central to mathematical problem solving, serving as essential building blocks for tackling increasingly complex challenges. Recent work, TroVE, argues that code-generating Large Language Models (LLMs) can benefit similarly on the MATH benchmark by inducing and reusing higher-level toolboxes. By allocating computational budget across an ensemble of three modes -- directly generating code, creating tools, and reusing tools -- TroVE claims to outperform a PRIMITIVE baseline that only performs direct generation. However, recent analysis (Berlot-Attwell et al., 2024) casts doubt on these gains, noting that the tools created are often trivial or rarely reused, suggesting that improvements may stem from self-consistency or self-correction. In this work, we re-evaluate TroVE on MATH, analyze the impact of each of its modes, and show that its benefit does not come from these mechanisms, but simply from a higher computational budget spent for TroVE compared to PRIMITIVE. To this end, we also perform a small correction in the original implementation of TroVE's selection mechanism, boosting TroVE's performance on MATH by 3\% in accuracy. After matching for compute, the benefit of TroVE reduces to a marginal improvement of 1\%, suggesting that this toolbox approach does not provide a significant benefit on MATH.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22069
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Compute-Matched Re-Evaluation of TroVE on MATH
Sesterhenn, Tobias
Berlot-Attwell, Ian
Zenkner, Janis
Bartelt, Christian
Programming Languages
Artificial Intelligence
Reusing established theorems and formulas is central to mathematical problem solving, serving as essential building blocks for tackling increasingly complex challenges. Recent work, TroVE, argues that code-generating Large Language Models (LLMs) can benefit similarly on the MATH benchmark by inducing and reusing higher-level toolboxes. By allocating computational budget across an ensemble of three modes -- directly generating code, creating tools, and reusing tools -- TroVE claims to outperform a PRIMITIVE baseline that only performs direct generation. However, recent analysis (Berlot-Attwell et al., 2024) casts doubt on these gains, noting that the tools created are often trivial or rarely reused, suggesting that improvements may stem from self-consistency or self-correction. In this work, we re-evaluate TroVE on MATH, analyze the impact of each of its modes, and show that its benefit does not come from these mechanisms, but simply from a higher computational budget spent for TroVE compared to PRIMITIVE. To this end, we also perform a small correction in the original implementation of TroVE's selection mechanism, boosting TroVE's performance on MATH by 3\% in accuracy. After matching for compute, the benefit of TroVE reduces to a marginal improvement of 1\%, suggesting that this toolbox approach does not provide a significant benefit on MATH.
title A Compute-Matched Re-Evaluation of TroVE on MATH
topic Programming Languages
Artificial Intelligence
url https://arxiv.org/abs/2507.22069