ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Raoyuan, Liu, Yihong, Du, Yupei, Schütze, Hinrich, Hedderich, Michael A.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910263960141824
author Zhao, Raoyuan
Liu, Yihong
Du, Yupei
Schütze, Hinrich
Hedderich, Michael A.
author_facet Zhao, Raoyuan
Liu, Yihong
Du, Yupei
Schütze, Hinrich
Hedderich, Michael A.
contents Mathematical reasoning benchmarks are vital for evaluating large language models (LLMs), but many are static and repeatedly exposed through public evaluation and training pipelines, making it difficult to separate genuine reasoning from memorization. Meanwhile, manually constructing new math problems with reliable answers remains costly. We introduce ReverseMath, a scalable method for generating new math problems through answer inversion. Given a problem and its answer, ReverseMath masks a numerical value in the original problem, treats the original answer as a known condition, and rewrites the problem so that the masked value becomes the new answer. The generated problem reverses the original input-output relation, making its answer known by construction. We study ReverseMath for both evaluation and training. For evaluation, paired original/reversed problems reveal substantial behavioral shifts: models sometimes fail on reversed problems and even incorrectly output the original answer, suggesting memorization-like behavior. For training, ReverseMath provides automatically labeled reversed problems as data augmentation for reinforcement learning (RL). Experiments show that including ReverseMath-generated data improves mathematical reasoning performance across multiple benchmarks, demonstrating its value as both an analysis tool and a scalable source of verifiable training data.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27709
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
Zhao, Raoyuan
Liu, Yihong
Du, Yupei
Schütze, Hinrich
Hedderich, Michael A.
Computation and Language
Mathematical reasoning benchmarks are vital for evaluating large language models (LLMs), but many are static and repeatedly exposed through public evaluation and training pipelines, making it difficult to separate genuine reasoning from memorization. Meanwhile, manually constructing new math problems with reliable answers remains costly. We introduce ReverseMath, a scalable method for generating new math problems through answer inversion. Given a problem and its answer, ReverseMath masks a numerical value in the original problem, treats the original answer as a known condition, and rewrites the problem so that the masked value becomes the new answer. The generated problem reverses the original input-output relation, making its answer known by construction. We study ReverseMath for both evaluation and training. For evaluation, paired original/reversed problems reveal substantial behavioral shifts: models sometimes fail on reversed problems and even incorrectly output the original answer, suggesting memorization-like behavior. For training, ReverseMath provides automatically labeled reversed problems as data augmentation for reinforcement learning (RL). Experiments show that including ReverseMath-generated data improves mathematical reasoning performance across multiple benchmarks, demonstrating its value as both an analysis tool and a scalable source of verifiable training data.
title ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
topic Computation and Language
url https://arxiv.org/abs/2605.27709