Unlocking Multimodal Mathematical Reasoning via Process Reward Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Ruilin, Zheng, Zhuofan, Wang, Yifan, Ni, Xinzhe, Lin, Zicheng, Jiang, Songtao, Yu, Yiyao, Shi, Chufan, Wang, Lei, Chu, Ruihang, Zeng, Jin, Yang, Yujiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
Hint-enhanced In-Context Learning wakes Large Language Models up for knowledge-intensive tasks
von: Wang, Yifan, et al.
Veröffentlicht: (2023)
von: Wang, Yifan, et al.
Veröffentlicht: (2023)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
Exploring the Mystery of Influential Data for Mathematical Reasoning
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
Generative Universal Verifier as Multimodal Meta-Reasoner
von: Zhang, Xinchen, et al.
Veröffentlicht: (2025)
von: Zhang, Xinchen, et al.
Veröffentlicht: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
ContextVis: Envision Contextual Learning and Interaction with Generative Models
von: Shui, Bo, et al.
Veröffentlicht: (2024)
von: Shui, Bo, et al.
Veröffentlicht: (2024)
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
von: Ni, Xinzhe, et al.
Veröffentlicht: (2022)
von: Ni, Xinzhe, et al.
Veröffentlicht: (2022)
MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence
von: Chen, Yifan, et al.
Veröffentlicht: (2026)
von: Chen, Yifan, et al.
Veröffentlicht: (2026)
From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation
von: Gu, Tianle, et al.
Veröffentlicht: (2026)
von: Gu, Tianle, et al.
Veröffentlicht: (2026)
LLM2: Let Large Language Models Harness System 2 Reasoning
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
A Thorough Examination of Decoding Methods in the Era of LLMs
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
Reward Modeling from Natural Language Human Feedback
von: Wang, Zongqi, et al.
Veröffentlicht: (2026)
von: Wang, Zongqi, et al.
Veröffentlicht: (2026)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
von: Wang, Tianqi, et al.
Veröffentlicht: (2024)
von: Wang, Tianqi, et al.
Veröffentlicht: (2024)
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
Incremental Residual Concept Bottleneck Models
von: Shang, Chenming, et al.
Veröffentlicht: (2024)
von: Shang, Chenming, et al.
Veröffentlicht: (2024)
Advances in Engineered Virus‐Like Particles for Applications in Nanomedicine
von: Qingxia Shi, et al.
Veröffentlicht: (2026)
von: Qingxia Shi, et al.
Veröffentlicht: (2026)
Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
von: Yu, Yiyao, et al.
Veröffentlicht: (2025)
von: Yu, Yiyao, et al.
Veröffentlicht: (2025)
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
von: Yang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Yang, Zhaohui, et al.
Veröffentlicht: (2025)
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
von: Ren, Yiming, et al.
Veröffentlicht: (2025)
von: Ren, Yiming, et al.
Veröffentlicht: (2025)
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
von: Hu, Yulan, et al.
Veröffentlicht: (2025)
von: Hu, Yulan, et al.
Veröffentlicht: (2025)
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
von: Chen, Zhengyu, et al.
Veröffentlicht: (2025)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
von: Li, Yu, et al.
Veröffentlicht: (2024)
von: Li, Yu, et al.
Veröffentlicht: (2024)
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
von: Chen, Yuqing, et al.
Veröffentlicht: (2025)
Velocity-Space 3D Asset Editing
von: Liu, Hao, et al.
Veröffentlicht: (2026)
von: Liu, Hao, et al.
Veröffentlicht: (2026)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning
von: Le, Yifan
Veröffentlicht: (2026)
von: Le, Yifan
Veröffentlicht: (2026)
ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
von: Xu, Rui, et al.
Veröffentlicht: (2025)
von: Xu, Rui, et al.
Veröffentlicht: (2025)
Asymmetric Idiosyncrasies in Multimodal Models
von: Tao, Muzi, et al.
Veröffentlicht: (2026)
von: Tao, Muzi, et al.
Veröffentlicht: (2026)
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
von: Zhuang, Wenwen, et al.
Veröffentlicht: (2024)
von: Zhuang, Wenwen, et al.
Veröffentlicht: (2024)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
von: Wang, Zhijie
Veröffentlicht: (2026)
von: Wang, Zhijie
Veröffentlicht: (2026)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
von: Cao, Qi, et al.
Veröffentlicht: (2025)
von: Cao, Qi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026) -
Hint-enhanced In-Context Learning wakes Large Language Models up for knowledge-intensive tasks
von: Wang, Yifan, et al.
Veröffentlicht: (2023) -
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
von: Lin, Zicheng, et al.
Veröffentlicht: (2024) -
Exploring the Mystery of Influential Data for Mathematical Reasoning
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024) -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
von: Ren, Yiming, et al.
Veröffentlicht: (2026)