Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aggarwal, Pranjal, Ghazvininejad, Marjan, Kim, Seungone, Kulikov, Ilia, Lanchantin, Jack, Li, Xian, Li, Tianjian, Liu, Bo, Neubig, Graham, Ovalle, Anaelia, Saha, Swarnadeep, Sukhbaatar, Sainbayar, Welleck, Sean, Weston, Jason, Whitehouse, Chenxi, Williams, Adina, Xu, Jing, Yu, Ping, Yuan, Weizhe, Zhang, Jingyu, Zhao, Wenting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
SPICE: Self-Play In Corpus Environments Improves Reasoning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
The Majority is not always right: RL training for solution aggregation
von: Zhao, Wenting, et al.
Veröffentlicht: (2025)
von: Zhao, Wenting, et al.
Veröffentlicht: (2025)
Adaptive Decoding via Latent Preference Optimization
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
Diverse Preference Optimization
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Bridging Offline and Online Reinforcement Learning for LLMs
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
Following Length Constraints in Instructions
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Gym-Anything: Turn any Software into an Agent Environment
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
Thinking LLMs: General Instruction Following with Thought Generation
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
Multi-Token Attention
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
Contextual Position Encoding: Learning to Count What's Important
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
von: Xu, Jing, et al.
Veröffentlicht: (2023)
von: Xu, Jing, et al.
Veröffentlicht: (2023)
Iterative Reasoning Preference Optimization
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
Jointly Reinforcing Diversity and Quality in Language Model Generations
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
R.I.P.: Better Models by Survival of the Fittest Prompts
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2024)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2024)
Self-Rewarding Language Models
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Agentic-R1: Distilled Dual-Strategy Reasoning
von: Du, Weihua, et al.
Veröffentlicht: (2025)
von: Du, Weihua, et al.
Veröffentlicht: (2025)
From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
von: Agarwal, Anmol, et al.
Veröffentlicht: (2026)
von: Agarwal, Anmol, et al.
Veröffentlicht: (2026)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
von: Tao, Leitian, et al.
Veröffentlicht: (2025)
von: Tao, Leitian, et al.
Veröffentlicht: (2025)
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
Distilling System 2 into System 1
von: Yu, Ping, et al.
Veröffentlicht: (2024)
von: Yu, Ping, et al.
Veröffentlicht: (2024)
Towards Knowledge-Grounded Natural Language Understanding and Generation
von: Whitehouse, Chenxi
Veröffentlicht: (2024)
von: Whitehouse, Chenxi
Veröffentlicht: (2024)
Training Large Language Models to Reason in a Continuous Latent Space
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
LLM Pretraining with Continuous Concepts
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2025)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025) -
SPICE: Self-Play In Corpus Environments Improves Reasoning
von: Liu, Bo, et al.
Veröffentlicht: (2025) -
The Majority is not always right: RL training for solution aggregation
von: Zhao, Wenting, et al.
Veröffentlicht: (2025) -
Adaptive Decoding via Latent Preference Optimization
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024) -
Diverse Preference Optimization
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)