Generative Verifiers: Reward Modeling as Next-Token Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lunjun, Hosseini, Arian, Bansal, Hritik, Kazemi, Mehran, Kumar, Aviral, Agarwal, Rishabh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
by: Setlur, Amrith, et al.
Published: (2024)
by: Setlur, Amrith, et al.
Published: (2024)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
by: Singhi, Nishad, et al.
Published: (2025)
by: Singhi, Nishad, et al.
Published: (2025)
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
by: Sareen, Kusha, et al.
Published: (2025)
by: Sareen, Kusha, et al.
Published: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
Not All LLM Reasoners Are Created Equal
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
CONFEX: Uncertainty-Aware Counterfactual Explanations with Conformal Guarantees
by: Bilkhoo, Aman, et al.
Published: (2025)
by: Bilkhoo, Aman, et al.
Published: (2025)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
by: Bansal, Hritik, et al.
Published: (2023)
by: Bansal, Hritik, et al.
Published: (2023)
Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
by: Chandra, Abhranil, et al.
Published: (2025)
by: Chandra, Abhranil, et al.
Published: (2025)
Verifiably Robust Conformal Prediction
by: Jeary, Linus, et al.
Published: (2024)
by: Jeary, Linus, et al.
Published: (2024)
Certified Guidance for Planning with Deep Generative Models
by: Giacomarra, Francesco, et al.
Published: (2025)
by: Giacomarra, Francesco, et al.
Published: (2025)
Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades
by: Pona, Edoardo, et al.
Published: (2026)
by: Pona, Edoardo, et al.
Published: (2026)
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
by: Zhang, Lunjun, et al.
Published: (2026)
by: Zhang, Lunjun, et al.
Published: (2026)
Trajeglish: Traffic Modeling as Next-Token Prediction
by: Philion, Jonah, et al.
Published: (2023)
by: Philion, Jonah, et al.
Published: (2023)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
by: Zhang, Kaiyi, et al.
Published: (2026)
by: Zhang, Kaiyi, et al.
Published: (2026)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
by: Hu, Haoyu, et al.
Published: (2026)
by: Hu, Haoyu, et al.
Published: (2026)
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
by: Huang, Guanhua, et al.
Published: (2025)
by: Huang, Guanhua, et al.
Published: (2025)
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
by: Kim, Yoonjeon, et al.
Published: (2025)
by: Kim, Yoonjeon, et al.
Published: (2025)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
by: Suvarna, Ashima, et al.
Published: (2026)
by: Suvarna, Ashima, et al.
Published: (2026)
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Trust, but Verify: Peeling Low-Bit Transformer Networks for Training Monitoring
by: Eamaz, Arian, et al.
Published: (2026)
by: Eamaz, Arian, et al.
Published: (2026)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
Provable Generalization in Overparameterized Neural Nets
by: Dhingra, Aviral
Published: (2025)
by: Dhingra, Aviral
Published: (2025)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
Towards Understanding the Universality of Transformers for Next-Token Prediction
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
Adaptively Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
by: Schneider, Johannes
Published: (2024)
by: Schneider, Johannes
Published: (2024)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
ENTP: Encoder-only Next Token Prediction
by: Ewer, Ethan, et al.
Published: (2024)
by: Ewer, Ethan, et al.
Published: (2024)
Prot2Token: A Unified Framework for Protein Modeling via Next-Token Prediction
by: Pourmirzaei, Mahdi, et al.
Published: (2025)
by: Pourmirzaei, Mahdi, et al.
Published: (2025)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
Scaling Next-Brain-Token Prediction for MEG
by: Csaky, Richard
Published: (2026)
by: Csaky, Richard
Published: (2026)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
by: Chow, Yinlam, et al.
Published: (2024)
by: Chow, Yinlam, et al.
Published: (2024)
A Law of Next-Token Prediction in Large Language Models
by: He, Hangfeng, et al.
Published: (2024)
by: He, Hangfeng, et al.
Published: (2024)
Differentially Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
Similar Items
-
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024) -
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
by: Setlur, Amrith, et al.
Published: (2024) -
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
by: Singhi, Nishad, et al.
Published: (2025) -
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
by: Sareen, Kusha, et al.
Published: (2025) -
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)