Evaluating Robustness of Reward Models for Mathematical Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Sunghwan, Kang, Dongjin, Kwon, Taeyoon, Chae, Hyungjoo, Won, Jungsoo, Lee, Dongha, Yeo, Jinyoung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
por: Kim, Sunghwan, et al.
Publicado: (2025)
por: Kim, Sunghwan, et al.
Publicado: (2025)
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
por: Ko, Sungho, et al.
Publicado: (2024)
por: Ko, Sungho, et al.
Publicado: (2024)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
por: Park, Jungsoo, et al.
Publicado: (2026)
por: Park, Jungsoo, et al.
Publicado: (2026)
Large Language Models are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales
por: Kwon, Taeyoon, et al.
Publicado: (2023)
por: Kwon, Taeyoon, et al.
Publicado: (2023)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
por: Seo, Yeongbin, et al.
Publicado: (2025)
por: Seo, Yeongbin, et al.
Publicado: (2025)
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
por: Ong, Kai Tzu-iunn, et al.
Publicado: (2024)
por: Ong, Kai Tzu-iunn, et al.
Publicado: (2024)
Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
por: Kang, Dongjin, et al.
Publicado: (2024)
por: Kang, Dongjin, et al.
Publicado: (2024)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
por: Kim, Seoyeon, et al.
Publicado: (2024)
por: Kim, Seoyeon, et al.
Publicado: (2024)
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
por: Kwak, Beong-woo, et al.
Publicado: (2025)
por: Kwak, Beong-woo, et al.
Publicado: (2025)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
por: Choi, Dongwook, et al.
Publicado: (2025)
por: Choi, Dongwook, et al.
Publicado: (2025)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
por: Kim, Sunghwan, et al.
Publicado: (2026)
por: Kim, Sunghwan, et al.
Publicado: (2026)
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
por: Chon, Heejae, et al.
Publicado: (2024)
por: Chon, Heejae, et al.
Publicado: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
por: Zhang, Zhenru, et al.
Publicado: (2025)
por: Zhang, Zhenru, et al.
Publicado: (2025)
Anticipatory Evaluation of Language Models
por: Park, Jungsoo, et al.
Publicado: (2025)
por: Park, Jungsoo, et al.
Publicado: (2025)
Unsupervised Robust Cross-Lingual Entity Alignment via Neighbor Triple Matching with Entity and Relation Texts
por: Yoon, Soojin, et al.
Publicado: (2024)
por: Yoon, Soojin, et al.
Publicado: (2024)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
por: Luo, Ruilin, et al.
Publicado: (2025)
por: Luo, Ruilin, et al.
Publicado: (2025)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
por: Prahlad, Deeksha, et al.
Publicado: (2025)
por: Prahlad, Deeksha, et al.
Publicado: (2025)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
por: Hwang, Yeonjun, et al.
Publicado: (2026)
por: Hwang, Yeonjun, et al.
Publicado: (2026)
On the Robustness of Reward Models for Language Model Alignment
por: Hong, Jiwoo, et al.
Publicado: (2025)
por: Hong, Jiwoo, et al.
Publicado: (2025)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
por: Oh, Jungwoo, et al.
Publicado: (2026)
por: Oh, Jungwoo, et al.
Publicado: (2026)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
por: Zheng, Congmin, et al.
Publicado: (2025)
por: Zheng, Congmin, et al.
Publicado: (2025)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
por: Lee, Seungbeen, et al.
Publicado: (2024)
por: Lee, Seungbeen, et al.
Publicado: (2024)
COCOA: CBT-based Conversational Counseling Agent using Memory Specialized in Cognitive Distortions and Dynamic Prompt
por: Lee, Suyeon, et al.
Publicado: (2024)
por: Lee, Suyeon, et al.
Publicado: (2024)
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
por: Sun, Mingzhong, et al.
Publicado: (2026)
por: Sun, Mingzhong, et al.
Publicado: (2026)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
por: Park, Jungsoo, et al.
Publicado: (2025)
por: Park, Jungsoo, et al.
Publicado: (2025)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
por: Park, Jinyoung, et al.
Publicado: (2023)
por: Park, Jinyoung, et al.
Publicado: (2023)
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
por: Kim, Sunghwan, et al.
Publicado: (2026)
por: Kim, Sunghwan, et al.
Publicado: (2026)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
por: Kang, Deokhyung, et al.
Publicado: (2025)
por: Kang, Deokhyung, et al.
Publicado: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
por: Hao, Yuren, et al.
Publicado: (2025)
por: Hao, Yuren, et al.
Publicado: (2025)
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement
por: Kim, Hana, et al.
Publicado: (2024)
por: Kim, Hana, et al.
Publicado: (2024)
Process Reward Models That Think
por: Khalifa, Muhammad, et al.
Publicado: (2025)
por: Khalifa, Muhammad, et al.
Publicado: (2025)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
por: Huang, Yuzhen, et al.
Publicado: (2025)
por: Huang, Yuzhen, et al.
Publicado: (2025)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
por: Moon, Seungjun, et al.
Publicado: (2023)
por: Moon, Seungjun, et al.
Publicado: (2023)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
por: Anugraha, David, et al.
Publicado: (2025)
por: Anugraha, David, et al.
Publicado: (2025)
Evaluating the Robustness of Analogical Reasoning in Large Language Models
por: Lewis, Martha, et al.
Publicado: (2024)
por: Lewis, Martha, et al.
Publicado: (2024)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
por: Li, Xiaoyuan, et al.
Publicado: (2024)
por: Li, Xiaoyuan, et al.
Publicado: (2024)
RM-R1: Reward Modeling as Reasoning
por: Chen, Xiusi, et al.
Publicado: (2025)
por: Chen, Xiusi, et al.
Publicado: (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
por: Gureja, Srishti, et al.
Publicado: (2024)
por: Gureja, Srishti, et al.
Publicado: (2024)
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
por: Chae, Hyungjoo, et al.
Publicado: (2024)
por: Chae, Hyungjoo, et al.
Publicado: (2024)
OvA-LP: A Simple and Efficient Framework for Federated Learning on Non-IID Data
por: Park, Dongjin, et al.
Publicado: (2025)
por: Park, Dongjin, et al.
Publicado: (2025)
Ejemplares similares
-
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
por: Kim, Sunghwan, et al.
Publicado: (2025) -
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
por: Ko, Sungho, et al.
Publicado: (2024) -
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
por: Park, Jungsoo, et al.
Publicado: (2026) -
Large Language Models are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales
por: Kwon, Taeyoon, et al.
Publicado: (2023) -
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
por: Seo, Yeongbin, et al.
Publicado: (2025)