Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Mahdavi, Sadegh, Kisacanin, Branislav, Toshniwal, Shubham, Du, Wei, Moshkov, Ivan, Armstrong, George, Liao, Renjie, Thrampoulidis, Christos, Gitman, Igor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
by: Du, Wei, et al.
Published: (2025)
by: Du, Wei, et al.
Published: (2025)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
by: Toshniwal, Shubham, et al.
Published: (2024)
by: Toshniwal, Shubham, et al.
Published: (2024)
GenSelect: A Generative Approach to Best-of-N
by: Toshniwal, Shubham, et al.
Published: (2025)
by: Toshniwal, Shubham, et al.
Published: (2025)
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023)
by: Mahdavi, Sadegh, et al.
Published: (2023)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
by: Toshniwal, Shubham, et al.
Published: (2024)
by: Toshniwal, Shubham, et al.
Published: (2024)
AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
by: Moshkov, Ivan, et al.
Published: (2025)
by: Moshkov, Ivan, et al.
Published: (2025)
Learning Generative Selection for Best-of-N
by: Toshniwal, Shubham, et al.
Published: (2026)
by: Toshniwal, Shubham, et al.
Published: (2026)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
The Challenge of Teaching Reasoning to LLMs Without RL or Distillation
by: Du, Wei, et al.
Published: (2025)
by: Du, Wei, et al.
Published: (2025)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
From Graph Diffusion to Graph Classification
by: Xian, Jia Jun Cheng, et al.
Published: (2024)
by: Xian, Jia Jun Cheng, et al.
Published: (2024)
Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining
by: Stromberg, Nathan, et al.
Published: (2025)
by: Stromberg, Nathan, et al.
Published: (2025)
Code Pretraining Improves Entity Tracking Abilities of Language Models
by: Kim, Najoung, et al.
Published: (2024)
by: Kim, Najoung, et al.
Published: (2024)
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
by: Vakilian, Vala, et al.
Published: (2025)
by: Vakilian, Vala, et al.
Published: (2025)
Memory capacity of two layer neural networks with smooth activations
by: Madden, Liam, et al.
Published: (2023)
by: Madden, Liam, et al.
Published: (2023)
Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features
by: Zhao, Yize, et al.
Published: (2026)
by: Zhao, Yize, et al.
Published: (2026)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
by: Zhao, Yize, et al.
Published: (2025)
by: Zhao, Yize, et al.
Published: (2025)
Supervised Contrastive Representation Learning: Landscape Analysis with Unconstrained Features
by: Behnia, Tina, et al.
Published: (2024)
by: Behnia, Tina, et al.
Published: (2024)
NeMo-Inspector: A Visualization Tool for LLM Generation Analysis
by: Gitman, Daria, et al.
Published: (2025)
by: Gitman, Daria, et al.
Published: (2025)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
by: Behnia, Tina, et al.
Published: (2025)
by: Behnia, Tina, et al.
Published: (2025)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
by: Zhao, Yize, et al.
Published: (2024)
by: Zhao, Yize, et al.
Published: (2024)
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
by: Deng, Wenlong, et al.
Published: (2023)
by: Deng, Wenlong, et al.
Published: (2023)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
by: Taheri, Hossein, et al.
Published: (2024)
by: Taheri, Hossein, et al.
Published: (2024)
Next-token prediction capacity: general upper bounds and a lower bound for transformers
by: Madden, Liam, et al.
Published: (2024)
by: Madden, Liam, et al.
Published: (2024)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Autograding Mathematical Induction Proofs with Natural Language Processing
by: Zhao, Chenyan, et al.
Published: (2024)
by: Zhao, Chenyan, et al.
Published: (2024)
Programs Versus Finite Tree-Programs
by: Moshkov, Mikhail
Published: (2025)
by: Moshkov, Mikhail
Published: (2025)
Algorithmic Problems for Computation Trees
by: Moshkov, Mikhail
Published: (2025)
by: Moshkov, Mikhail
Published: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
by: Naik, Aaditya, et al.
Published: (2026)
by: Naik, Aaditya, et al.
Published: (2026)
Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
by: Lifshitz, Shalev, et al.
Published: (2025)
by: Lifshitz, Shalev, et al.
Published: (2025)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
by: Garrod, Connall, et al.
Published: (2025)
by: Garrod, Connall, et al.
Published: (2025)
Leveraging Environment Interaction for Automated PDDL Translation and Planning with Large Language Models
by: Mahdavi, Sadegh, et al.
Published: (2024)
by: Mahdavi, Sadegh, et al.
Published: (2024)
LeanTutor: Towards a Verified AI Mathematical Proof Tutor
by: Patel, Manooshree, et al.
Published: (2025)
by: Patel, Manooshree, et al.
Published: (2025)
LeanTutor: Towards a Verified AI Mathematical Proof Tutor
by: Patel, Manooshree, et al.
Published: (2026)
by: Patel, Manooshree, et al.
Published: (2026)
ProofWright: Towards Agentic Formal Verification of CUDA
by: Chatterjee, Bodhisatwa, et al.
Published: (2025)
by: Chatterjee, Bodhisatwa, et al.
Published: (2025)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025)
by: Deora, Puneesh, et al.
Published: (2025)
Similar Items
-
Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
by: Du, Wei, et al.
Published: (2025) -
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
by: Toshniwal, Shubham, et al.
Published: (2024) -
GenSelect: A Generative Approach to Best-of-N
by: Toshniwal, Shubham, et al.
Published: (2025) -
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023) -
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
by: Toshniwal, Shubham, et al.
Published: (2024)