Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yuan, Youliang, Mang, Qiuyang, Chen, Jingbang, Wan, Hong, Liu, Xiaoyuan, Xu, Junjielong, Huang, Jen-tse, Wang, Wenxuan, Jiao, Wenxiang, He, Pinjia |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
par: Wan, Yuxuan, et autres
Publié: (2024)
par: Wan, Yuxuan, et autres
Publié: (2024)
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
par: Yuan, Youliang, et autres
Publié: (2023)
par: Yuan, Youliang, et autres
Publié: (2023)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
par: Yuan, Youliang, et autres
Publié: (2024)
par: Yuan, Youliang, et autres
Publié: (2024)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
par: Yuan, Youliang, et autres
Publié: (2025)
par: Yuan, Youliang, et autres
Publié: (2025)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
par: Liu, Xiaoyuan, et autres
Publié: (2024)
par: Liu, Xiaoyuan, et autres
Publié: (2024)
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
par: Wang, Wenxuan, et autres
Publié: (2024)
par: Wang, Wenxuan, et autres
Publié: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
par: Wang, Wenxuan, et autres
Publié: (2025)
par: Wang, Wenxuan, et autres
Publié: (2025)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
par: Huang, Jen-Tse, et autres
Publié: (2025)
par: Huang, Jen-Tse, et autres
Publié: (2025)
All Languages Matter: On the Multilingual Safety of Large Language Models
par: Wang, Wenxuan, et autres
Publié: (2023)
par: Wang, Wenxuan, et autres
Publié: (2023)
Scalable Supervising Software Agents with Patch Reasoner
par: Xu, Junjielong, et autres
Publié: (2025)
par: Xu, Junjielong, et autres
Publié: (2025)
Step-wise Rubric Rewards for LLM Reasoning
par: Xie, Weichu, et autres
Publié: (2026)
par: Xie, Weichu, et autres
Publié: (2026)
On the Shortcut Learning in Multilingual Neural Machine Translation
par: Wang, Wenxuan, et autres
Publié: (2024)
par: Wang, Wenxuan, et autres
Publié: (2024)
Learning to Ask: When LLM Agents Meet Unclear Instruction
par: Wang, Wenxuan, et autres
Publié: (2024)
par: Wang, Wenxuan, et autres
Publié: (2024)
Aligning the Objective of LLM-based Program Repair
par: Xu, Junjielong, et autres
Publié: (2024)
par: Xu, Junjielong, et autres
Publié: (2024)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
par: Huang, Jen-tse, et autres
Publié: (2025)
par: Huang, Jen-tse, et autres
Publié: (2025)
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench
par: Huang, Jen-tse, et autres
Publié: (2023)
par: Huang, Jen-tse, et autres
Publié: (2023)
Nearly Optimal Internal Dictionary Matching
par: Chen, Jingbang, et autres
Publié: (2023)
par: Chen, Jingbang, et autres
Publié: (2023)
AL-Bench: A Benchmark for Automatic Logging
par: Tan, Boyin, et autres
Publié: (2025)
par: Tan, Boyin, et autres
Publié: (2025)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
par: Wang, Wenxuan, et autres
Publié: (2024)
par: Wang, Wenxuan, et autres
Publié: (2024)
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
par: Wang, Wenxuan, et autres
Publié: (2023)
par: Wang, Wenxuan, et autres
Publié: (2023)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
par: Huang, Jen-tse, et autres
Publié: (2024)
par: Huang, Jen-tse, et autres
Publié: (2024)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
par: Wang, Wenxuan, et autres
Publié: (2024)
par: Wang, Wenxuan, et autres
Publié: (2024)
Revisiting the Reliability of Psychological Scales on Large Language Models
par: Huang, Jen-tse, et autres
Publié: (2023)
par: Huang, Jen-tse, et autres
Publié: (2023)
AI Sees Your Location, But With A Bias Toward The Wealthy World
par: Huang, Jingyuan, et autres
Publié: (2025)
par: Huang, Jingyuan, et autres
Publié: (2025)
On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
par: Huang, Jen-tse, et autres
Publié: (2024)
par: Huang, Jen-tse, et autres
Publié: (2024)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
par: Zhao, Sihang, et autres
Publié: (2024)
par: Zhao, Sihang, et autres
Publié: (2024)
Scalable Algorithm for Finding Balanced Subgraphs with Tolerance in Signed Networks
par: Chen, Jingbang, et autres
Publié: (2024)
par: Chen, Jingbang, et autres
Publié: (2024)
Prompting for Automatic Log Template Extraction
par: Xu, Junjielong, et autres
Publié: (2023)
par: Xu, Junjielong, et autres
Publié: (2023)
On the Failure of Latent State Persistence in Large Language Models
par: Huang, Jen-tse, et autres
Publié: (2025)
par: Huang, Jen-tse, et autres
Publié: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
par: Liu, Xiaoyuan, et autres
Publié: (2025)
par: Liu, Xiaoyuan, et autres
Publié: (2025)
Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
par: Li, Changyue, et autres
Publié: (2025)
par: Li, Changyue, et autres
Publié: (2025)
SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
par: Zhao, Sihang, et autres
Publié: (2026)
par: Zhao, Sihang, et autres
Publié: (2026)
CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
par: Lam, Man Ho, et autres
Publié: (2025)
par: Lam, Man Ho, et autres
Publié: (2025)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
par: Huang, Jen-tse, et autres
Publié: (2023)
par: Huang, Jen-tse, et autres
Publié: (2023)
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
par: Liu, Tianci, et autres
Publié: (2025)
par: Liu, Tianci, et autres
Publié: (2025)
A Goal-Driven Survey on Root Cause Analysis
par: Fang, Aoyang, et autres
Publié: (2025)
par: Fang, Aoyang, et autres
Publié: (2025)
SWE-Manager: Selecting and Synthesizing Golden Proposals Before Coding
par: Tan, Boyin, et autres
Publié: (2026)
par: Tan, Boyin, et autres
Publié: (2026)
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
par: Jia, Mengzhao, et autres
Publié: (2025)
par: Jia, Mengzhao, et autres
Publié: (2025)
MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context Learning
par: Yu, Jianbo, et autres
Publié: (2026)
par: Yu, Jianbo, et autres
Publié: (2026)
Scalable Approximate Biclique Counting over Large Bipartite Graphs
par: Chen, Jingbang, et autres
Publié: (2025)
par: Chen, Jingbang, et autres
Publié: (2025)
Documents similaires
-
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
par: Wan, Yuxuan, et autres
Publié: (2024) -
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
par: Yuan, Youliang, et autres
Publié: (2023) -
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
par: Yuan, Youliang, et autres
Publié: (2024) -
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
par: Yuan, Youliang, et autres
Publié: (2025) -
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
par: Liu, Xiaoyuan, et autres
Publié: (2024)