Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Sanders, Kate, Weir, Nathaniel, Chaudhary, Sapana, Bostrom, Kaj, Rangwala, Huzefa |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
by: Ou, Jiefu, et al.
Published: (2026)
by: Ou, Jiefu, et al.
Published: (2026)
Offline Learning and Forgetting for Reasoning with Large Language Models
by: Ni, Tianwei, et al.
Published: (2025)
by: Ni, Tianwei, et al.
Published: (2025)
ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models
by: He, Andre, et al.
Published: (2026)
by: He, Andre, et al.
Published: (2026)
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
by: Yang, Ke, et al.
Published: (2024)
by: Yang, Ke, et al.
Published: (2024)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
by: Ding, Liang
Published: (2026)
by: Ding, Liang
Published: (2026)
Risk-Averse Finetuning of Large Language Models
by: Chaudhary, Sapana, et al.
Published: (2025)
by: Chaudhary, Sapana, et al.
Published: (2025)
Learning to Reason via Program Generation, Emulation, and Search
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
by: Gunjal, Anisha, et al.
Published: (2025)
by: Gunjal, Anisha, et al.
Published: (2025)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Rubric-Guided Process Reward for Stepwise Model Routing
by: Ye, Shenghao, et al.
Published: (2026)
by: Ye, Shenghao, et al.
Published: (2026)
VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning
by: Singh, Vikash, et al.
Published: (2026)
by: Singh, Vikash, et al.
Published: (2026)
Reframing Tax Law Entailment as Analogical Reasoning
by: Zou, Xinrui, et al.
Published: (2024)
by: Zou, Xinrui, et al.
Published: (2024)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
by: Xu, Tianze, et al.
Published: (2026)
by: Xu, Tianze, et al.
Published: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
by: Ye, Zhiling, et al.
Published: (2025)
by: Ye, Zhiling, et al.
Published: (2025)
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
by: Fan, Zhiting, et al.
Published: (2026)
by: Fan, Zhiting, et al.
Published: (2026)
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
by: Xu, Yifei, et al.
Published: (2025)
by: Xu, Yifei, et al.
Published: (2025)
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
by: Sanders, Kate, et al.
Published: (2025)
by: Sanders, Kate, et al.
Published: (2025)
ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models
by: Chen, Bin, et al.
Published: (2025)
by: Chen, Bin, et al.
Published: (2025)
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
by: Lee, Jinu, et al.
Published: (2025)
by: Lee, Jinu, et al.
Published: (2025)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
by: Gao, Jie, et al.
Published: (2026)
by: Gao, Jie, et al.
Published: (2026)
Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning
by: Wang, Jingjing, et al.
Published: (2026)
by: Wang, Jingjing, et al.
Published: (2026)
The Art of Efficient Reasoning: Data, Reward, and Optimization
by: Wu, Taiqiang, et al.
Published: (2026)
by: Wu, Taiqiang, et al.
Published: (2026)
Configurable Preference Tuning with Rubric-Guided Synthetic Data
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
by: Zhao, Chengshuai, et al.
Published: (2026)
by: Zhao, Chengshuai, et al.
Published: (2026)
Exploring Reasoning Reward Model for Agents
by: Fan, Kaixuan, et al.
Published: (2026)
by: Fan, Kaixuan, et al.
Published: (2026)
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
by: Zhang, Yunyi, et al.
Published: (2026)
by: Zhang, Yunyi, et al.
Published: (2026)
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
by: Sun, Zhouhao, et al.
Published: (2026)
by: Sun, Zhouhao, et al.
Published: (2026)
GRAM: A Generative Foundation Reward Model for Reward Generalization
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation
by: Xu, Shanshan, et al.
Published: (2026)
by: Xu, Shanshan, et al.
Published: (2026)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL
by: Hua, Harper, et al.
Published: (2026)
by: Hua, Harper, et al.
Published: (2026)
Hierarchical Lexical Graph for Enhanced Multi-Hop Retrieval
by: Ghassel, Abdellah, et al.
Published: (2025)
by: Ghassel, Abdellah, et al.
Published: (2025)
Reasoning-Driven Synthetic Data Generation and Evaluation
by: Davidson, Tim R., et al.
Published: (2026)
by: Davidson, Tim R., et al.
Published: (2026)
Similar Items
-
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks
by: Feng, Yu, et al.
Published: (2025) -
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
by: Ou, Jiefu, et al.
Published: (2026) -
Offline Learning and Forgetting for Reasoning with Large Language Models
by: Ni, Tianwei, et al.
Published: (2025) -
ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models
by: He, Andre, et al.
Published: (2026) -
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
by: Yang, Ke, et al.
Published: (2024)