MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
Fuente:
arXiv
Saved in:
| Main Authors: | Mayilvahanan, Prasanna, Dominguez-Olmedo, Ricardo, Wiedemer, Thaddäus, Brendel, Wieland |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
by: Mayilvahanan, Prasanna, et al.
Published: (2025)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
by: Zeller, Jana, et al.
Published: (2026)
by: Zeller, Jana, et al.
Published: (2026)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
Provable Compositional Generalization for Object-Centric Learning
by: Wiedemer, Thaddäus, et al.
Published: (2023)
by: Wiedemer, Thaddäus, et al.
Published: (2023)
In Search of Forgotten Domain Generalization
by: Mayilvahanan, Prasanna, et al.
Published: (2024)
by: Mayilvahanan, Prasanna, et al.
Published: (2024)
VGGSounder: Audio-Visual Evaluations for Foundation Models
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
by: Yu, Zhouliang, et al.
Published: (2025)
by: Yu, Zhouliang, et al.
Published: (2025)
Position: An Empirically Grounded Identifiability Theory Will Accelerate Self-Supervised Learning Research
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Estimating Treatment Effects with Independent Component Analysis
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
Train-before-Test Harmonizes Language Model Rankings
by: Zhang, Guanhua, et al.
Published: (2025)
by: Zhang, Guanhua, et al.
Published: (2025)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
by: Yao, Jian, et al.
Published: (2025)
by: Yao, Jian, et al.
Published: (2025)
OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling
by: Lu, Hongliang, et al.
Published: (2025)
by: Lu, Hongliang, et al.
Published: (2025)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
by: Huang, Kaixuan, et al.
Published: (2025)
by: Huang, Kaixuan, et al.
Published: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
by: Puri, Isha, et al.
Published: (2026)
by: Puri, Isha, et al.
Published: (2026)
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
by: Ye, Chenlu, et al.
Published: (2025)
by: Ye, Chenlu, et al.
Published: (2025)
PANDA: Expanded Width-Aware Message Passing Beyond Rewiring
by: Choi, Jeongwhan, et al.
Published: (2024)
by: Choi, Jeongwhan, et al.
Published: (2024)
Cross-Entropy Is All You Need To Invert the Data Generating Process
by: Reizinger, Patrik, et al.
Published: (2024)
by: Reizinger, Patrik, et al.
Published: (2024)
The Evaluation Game: Beyond Static LLM Benchmarking
by: Wang, Paul, et al.
Published: (2026)
by: Wang, Paul, et al.
Published: (2026)
Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization
by: Wang, Dayu, et al.
Published: (2026)
by: Wang, Dayu, et al.
Published: (2026)
Computational Arbitrage in AI Model Markets
by: Olmedo, Ricardo, et al.
Published: (2026)
by: Olmedo, Ricardo, et al.
Published: (2026)
Video models are zero-shot learners and reasoners
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
by: Chen, Fan, et al.
Published: (2022)
by: Chen, Fan, et al.
Published: (2022)
Training on the Test Task Confounds Evaluation and Emergence
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
Spatio-Temporal Graphs Beyond Grids: Benchmark for Maritime Anomaly Detection
by: Kim, Jeehong, et al.
Published: (2025)
by: Kim, Jeehong, et al.
Published: (2025)
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
by: Yuan, Peiwen, et al.
Published: (2025)
by: Yuan, Peiwen, et al.
Published: (2025)
Revisiting Synthetic Human Trajectories: Imitative Generation and Benchmarks Beyond Datasaurus
by: Deng, Bangchao, et al.
Published: (2024)
by: Deng, Bangchao, et al.
Published: (2024)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Beyond Benchmarks: On The False Promise of AI Regulation
by: Stanovsky, Gabriel, et al.
Published: (2025)
by: Stanovsky, Gabriel, et al.
Published: (2025)
Beyond Isolated Clients: Integrating Graph-Based Embeddings into Event Sequence Models
by: Proshian, Harry, et al.
Published: (2026)
by: Proshian, Harry, et al.
Published: (2026)
Beyond Model Adaptation at Test Time: A Survey
by: Xiao, Zehao, et al.
Published: (2024)
by: Xiao, Zehao, et al.
Published: (2024)
Meta-World+: An Improved, Standardized, RL Benchmark
by: McLean, Reginald, et al.
Published: (2025)
by: McLean, Reginald, et al.
Published: (2025)
OGBench: Benchmarking Offline Goal-Conditioned RL
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
Beyond Fixed Variables: Expanding-variate Time Series Forecasting via Flat Scheme and Spatio-temporal Focal Learning
by: Ma, Minbo, et al.
Published: (2025)
by: Ma, Minbo, et al.
Published: (2025)
Chimera: State Space Models Beyond Sequences
by: Lahoti, Aakash, et al.
Published: (2025)
by: Lahoti, Aakash, et al.
Published: (2025)
Vanishing Feature: Diagnosing Model Merging and Beyond
by: Qu, Xingyu, et al.
Published: (2024)
by: Qu, Xingyu, et al.
Published: (2024)
Similar Items
-
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
by: Mayilvahanan, Prasanna, et al.
Published: (2025) -
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
by: Mayilvahanan, Prasanna, et al.
Published: (2023) -
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
by: Zeller, Jana, et al.
Published: (2026) -
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
by: Wiedemer, Thaddäus, et al.
Published: (2025) -
Provable Compositional Generalization for Object-Centric Learning
by: Wiedemer, Thaddäus, et al.
Published: (2023)