PRISM: Pushing the Frontier of Deep Think via Process Reward Model-Guided Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Sharma, Rituraj, Chen, Weiyuan, Provenzano, Noah, Vu, Tu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvoSkill: Automated Skill Discovery for Multi-Agent Systems
by: Alzubi, Salaheddin, et al.
Published: (2026)
by: Alzubi, Salaheddin, et al.
Published: (2026)
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
by: Balasubramanian, Rishab, et al.
Published: (2026)
by: Balasubramanian, Rishab, et al.
Published: (2026)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
PRISM: Parallel Reward Integration with Symmetry for MORL
by: van der Knaap, Finn, et al.
Published: (2026)
by: van der Knaap, Finn, et al.
Published: (2026)
SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
by: Pham, Thinh, et al.
Published: (2025)
by: Pham, Thinh, et al.
Published: (2025)
Rubric-Guided Process Reward for Stepwise Model Routing
by: Ye, Shenghao, et al.
Published: (2026)
by: Ye, Shenghao, et al.
Published: (2026)
PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models
by: Wang, Xuliang, et al.
Published: (2026)
by: Wang, Xuliang, et al.
Published: (2026)
ConPoSe: LLM-Guided Contact Point Selection for Scalable Cooperative Object Pushing
by: Steinkrüger, Noah, et al.
Published: (2025)
by: Steinkrüger, Noah, et al.
Published: (2025)
PRISM: Distributed Inference for Foundation Models at Edge
by: Qazi, Muhammad Azlan, et al.
Published: (2025)
by: Qazi, Muhammad Azlan, et al.
Published: (2025)
AEGIS: From Clues to Verdicts -- Graph-Guided Deep Vulnerability Reasoning via Dialectics and Meta-Auditing
by: Fang, Sen, et al.
Published: (2026)
by: Fang, Sen, et al.
Published: (2026)
Pushing the Frontier on Approximate EFX Allocations
by: Amanatidis, Georgios, et al.
Published: (2024)
by: Amanatidis, Georgios, et al.
Published: (2024)
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
by: Gan, Siyuan, et al.
Published: (2026)
by: Gan, Siyuan, et al.
Published: (2026)
PRISM: Projection-based Reward Integration for Scene-Aware Real-to-Sim-to-Real Transfer with Few Demonstrations
by: Sun, Haowen, et al.
Published: (2025)
by: Sun, Haowen, et al.
Published: (2025)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
by: Yang, Zeyu, et al.
Published: (2025)
by: Yang, Zeyu, et al.
Published: (2025)
GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
by: Falcon LLM Team, et al.
Published: (2026)
by: Falcon LLM Team, et al.
Published: (2026)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
by: Pala, Tej Deep, et al.
Published: (2025)
by: Pala, Tej Deep, et al.
Published: (2025)
Efficient Process Reward Model Training via Active Learning
by: Duan, Keyu, et al.
Published: (2025)
by: Duan, Keyu, et al.
Published: (2025)
WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
by: Yao, Yihang, et al.
Published: (2026)
by: Yao, Yihang, et al.
Published: (2026)
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
by: Lee, Kwanhee, et al.
Published: (2025)
by: Lee, Kwanhee, et al.
Published: (2025)
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
Pushing the boundary on Natural Language Inference
by: Miralles-González, Pablo, et al.
Published: (2025)
by: Miralles-González, Pablo, et al.
Published: (2025)
SmartSearch: Process Reward-Guided Query Refinement for Search Agents
by: Wen, Tongyu, et al.
Published: (2026)
by: Wen, Tongyu, et al.
Published: (2026)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
by: Rezaei, Mohammad, et al.
Published: (2026)
by: Rezaei, Mohammad, et al.
Published: (2026)
Visual Prompt Guided Unified Pushing Policy
by: Bui, Hieu, et al.
Published: (2026)
by: Bui, Hieu, et al.
Published: (2026)
Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review
by: Uehara, Masatoshi, et al.
Published: (2025)
by: Uehara, Masatoshi, et al.
Published: (2025)
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
by: Fan, Run-Ze, et al.
Published: (2025)
by: Fan, Run-Ze, et al.
Published: (2025)
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
by: Shinoda, Kazutoshi, et al.
Published: (2026)
by: Shinoda, Kazutoshi, et al.
Published: (2026)
Controllable and Verifiable Process Data Synthesis for Process Reward Models
by: Chi, Yinghui, et al.
Published: (2026)
by: Chi, Yinghui, et al.
Published: (2026)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026)
by: Ma, Rachel, et al.
Published: (2026)
How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models
by: Madur, Sai Sourabh
Published: (2026)
by: Madur, Sai Sourabh
Published: (2026)
ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems
by: Alzu'bi, Salaheddin, et al.
Published: (2026)
by: Alzu'bi, Salaheddin, et al.
Published: (2026)
PRISM: Privacy-preserving Inference System with Homomorphic Encryption and Modular Activation
by: Elkhatib, Zeinab, et al.
Published: (2025)
by: Elkhatib, Zeinab, et al.
Published: (2025)
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
by: Xie, Shaoan, et al.
Published: (2025)
by: Xie, Shaoan, et al.
Published: (2025)
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
by: Kim, Kwanyoung, et al.
Published: (2025)
by: Kim, Kwanyoung, et al.
Published: (2025)
Pushing the Limits of BFP on Narrow Precision LLM Inference
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
PRISM: Structured Optimization via Anisotropic Spectral Shaping
by: Yang, Yujie
Published: (2026)
by: Yang, Yujie
Published: (2026)
Similar Items
-
EvoSkill: Automated Skill Discovery for Multi-Agent Systems
by: Alzubi, Salaheddin, et al.
Published: (2026) -
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
by: Balasubramanian, Rishab, et al.
Published: (2026) -
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025) -
PRISM: Parallel Reward Integration with Symmetry for MORL
by: van der Knaap, Finn, et al.
Published: (2026) -
SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
by: Pham, Thinh, et al.
Published: (2025)