DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Qi, Xie, Pengtao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025)
by: Cao, Qi, et al.
Published: (2025)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
by: Zhang, Ruiyi, et al.
Published: (2025)
by: Zhang, Ruiyi, et al.
Published: (2025)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
by: Zhang, Ruiyi, et al.
Published: (2026)
by: Zhang, Ruiyi, et al.
Published: (2026)
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
by: Zeng, Thomas, et al.
Published: (2025)
by: Zeng, Thomas, et al.
Published: (2025)
EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing
by: Shihab, Ibne Farabi, et al.
Published: (2026)
by: Shihab, Ibne Farabi, et al.
Published: (2026)
ATLAS: Agentic Test-time Learning-to-Allocate Scaling
by: Qin, Peijia, et al.
Published: (2026)
by: Qin, Peijia, et al.
Published: (2026)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025)
by: Luo, Ruilin, et al.
Published: (2025)
Training Data Efficiency in Multimodal Process Reward Models
by: Li, Jinyuan, et al.
Published: (2026)
by: Li, Jinyuan, et al.
Published: (2026)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
by: Qin, Peijia, et al.
Published: (2026)
by: Qin, Peijia, et al.
Published: (2026)
Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
by: Cao, Qi, et al.
Published: (2026)
by: Cao, Qi, et al.
Published: (2026)
Adversarial Training for Process Reward Models
by: Juneja, Gurusha, et al.
Published: (2025)
by: Juneja, Gurusha, et al.
Published: (2025)
LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling
by: Cao, Qi, et al.
Published: (2026)
by: Cao, Qi, et al.
Published: (2026)
NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
by: Zhang, Zhi, et al.
Published: (2025)
by: Zhang, Zhi, et al.
Published: (2025)
Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense
by: Zhang, Shuhao, et al.
Published: (2026)
by: Zhang, Shuhao, et al.
Published: (2026)
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning
by: Mathew, Sheryl, et al.
Published: (2025)
by: Mathew, Sheryl, et al.
Published: (2025)
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
by: Feng, Zhangying, et al.
Published: (2025)
by: Feng, Zhangying, et al.
Published: (2025)
DreamGen: Unlocking Generalization in Robot Learning through Video World Models
by: Jang, Joel, et al.
Published: (2025)
by: Jang, Joel, et al.
Published: (2025)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
by: Wang, Weiyun, et al.
Published: (2025)
by: Wang, Weiyun, et al.
Published: (2025)
Efficient Process Reward Model Training via Active Learning
by: Duan, Keyu, et al.
Published: (2025)
by: Duan, Keyu, et al.
Published: (2025)
Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
by: He, Tao, et al.
Published: (2025)
by: He, Tao, et al.
Published: (2025)
A Foundational Multi-Modal Model for Few-Shot Learning
by: Dang, Pengtao, et al.
Published: (2025)
by: Dang, Pengtao, et al.
Published: (2025)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
by: Yang, Shidong, et al.
Published: (2026)
by: Yang, Shidong, et al.
Published: (2026)
BiLoRA: A Bi-level Optimization Framework for Overfitting-Resilient Low-Rank Adaptation of Large Pre-trained Models
by: Qiang, Rushi, et al.
Published: (2024)
by: Qiang, Rushi, et al.
Published: (2024)
Adversarial Training of Reward Models
by: Bukharin, Alexander, et al.
Published: (2025)
by: Bukharin, Alexander, et al.
Published: (2025)
Unlocking the Potential of Linear Networks for Irregular Multivariate Time Series Forecasting
by: Wang, Chengsen, et al.
Published: (2025)
by: Wang, Chengsen, et al.
Published: (2025)
Generative Modeling with Multi-Instance Reward Learning for E-commerce Creative Optimization
by: Gu, Qiaolei, et al.
Published: (2025)
by: Gu, Qiaolei, et al.
Published: (2025)
Unlocking the Potential of Model Calibration in Federated Learning
by: Chu, Yun-Wei, et al.
Published: (2024)
by: Chu, Yun-Wei, et al.
Published: (2024)
Unsupervised Process Reward Models
by: Gadetsky, Artyom, et al.
Published: (2026)
by: Gadetsky, Artyom, et al.
Published: (2026)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
by: Lee, Vint, et al.
Published: (2023)
by: Lee, Vint, et al.
Published: (2023)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
TapWeight: Reweighting Pretraining Objectives for Task-Adaptive Pretraining
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
Unlocking Potential Binders: Multimodal Pretraining DEL-Fusion for Denoising DNA-Encoded Libraries
by: Gu, Chunbin, et al.
Published: (2024)
by: Gu, Chunbin, et al.
Published: (2024)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
BiDoRA: Bi-level Optimization-Based Weight-Decomposed Low-Rank Adaptation
by: Qin, Peijia, et al.
Published: (2024)
by: Qin, Peijia, et al.
Published: (2024)
DreamLLM: Synergistic Multimodal Comprehension and Creation
by: Dong, Runpei, et al.
Published: (2023)
by: Dong, Runpei, et al.
Published: (2023)
Unlocking the Pre-Trained Model as a Dual-Alignment Calibrator for Post-Trained LLMs
by: Luo, Beier, et al.
Published: (2026)
by: Luo, Beier, et al.
Published: (2026)
Reward-Free Curricula for Training Robust World Models
by: Rigter, Marc, et al.
Published: (2023)
by: Rigter, Marc, et al.
Published: (2023)
Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
by: Lou, Xingzhou, et al.
Published: (2024)
by: Lou, Xingzhou, et al.
Published: (2024)
Similar Items
-
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025) -
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
by: Zhang, Ruiyi, et al.
Published: (2025) -
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
by: Zhang, Ruiyi, et al.
Published: (2026) -
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
by: Zeng, Thomas, et al.
Published: (2025) -
EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing
by: Shihab, Ibne Farabi, et al.
Published: (2026)