A Rubric-Supervised Critic from Sparse Real-World Outcomes
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Xingyao, Chen, Valerie, Ji, Heng, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Future-as-Label: Scalable Supervision from Real-World Outcomes
par: Turtel, Benjamin, et autres
Publié: (2026)
par: Turtel, Benjamin, et autres
Publié: (2026)
Gym-Anything: Turn any Software into an Agent Environment
par: Aggarwal, Pranjal, et autres
Publié: (2026)
par: Aggarwal, Pranjal, et autres
Publié: (2026)
Training Proactive and Personalized LLM Agents
par: Sun, Weiwei, et autres
Publié: (2025)
par: Sun, Weiwei, et autres
Publié: (2025)
CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets
par: Yuan, Lifan, et autres
Publié: (2023)
par: Yuan, Lifan, et autres
Publié: (2023)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
par: Xie, Lipeng, et autres
Publié: (2025)
par: Xie, Lipeng, et autres
Publié: (2025)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
par: Wang, Xingyao, et autres
Publié: (2023)
par: Wang, Xingyao, et autres
Publié: (2023)
Rubric-based On-policy Distillation
par: Fang, Junfeng, et autres
Publié: (2026)
par: Fang, Junfeng, et autres
Publié: (2026)
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards
par: Han, Zhixin, et autres
Publié: (2026)
par: Han, Zhixin, et autres
Publié: (2026)
Sparsely Supervised Diffusion
par: Zhao, Wenshuai, et autres
Publié: (2026)
par: Zhao, Wenshuai, et autres
Publié: (2026)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
par: Sharma, Manasi, et autres
Publié: (2025)
par: Sharma, Manasi, et autres
Publié: (2025)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
par: Wu, Peilin, et autres
Publié: (2026)
par: Wu, Peilin, et autres
Publié: (2026)
Bridging Policy and Real-World Dynamics: LLM-Augmented Rebalancing for Shared Micromobility Systems
par: Tan, Heng, et autres
Publié: (2026)
par: Tan, Heng, et autres
Publié: (2026)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
par: Qian, Cheng, et autres
Publié: (2025)
par: Qian, Cheng, et autres
Publié: (2025)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
par: Sheng, Leheng, et autres
Publié: (2026)
par: Sheng, Leheng, et autres
Publié: (2026)
Online Rubrics Elicitation from Pairwise Comparisons
par: Rezaei, MohammadHossein, et autres
Publié: (2025)
par: Rezaei, MohammadHossein, et autres
Publié: (2025)
Reinforcement Learning with Rubric Anchors
par: Huang, Zenan, et autres
Publié: (2025)
par: Huang, Zenan, et autres
Publié: (2025)
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
par: Nagar, Aishik, et autres
Publié: (2026)
par: Nagar, Aishik, et autres
Publié: (2026)
TOM-SWE: User Mental Modeling For Software Engineering Agents
par: Zhou, Xuhui, et autres
Publié: (2025)
par: Zhou, Xuhui, et autres
Publié: (2025)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
par: Liu, Dengcan, et autres
Publié: (2026)
par: Liu, Dengcan, et autres
Publié: (2026)
Robustness of Graph Self-Supervised Learning to Real-World Noise: A Case Study on Text-Driven Biomedical Graphs
par: Kabal, Othmane, et autres
Publié: (2026)
par: Kabal, Othmane, et autres
Publié: (2026)
Uncertainty-Calibrated Spatiotemporal Field Diffusion with Sparse Supervision
par: Valencia, Kevin, et autres
Publié: (2026)
par: Valencia, Kevin, et autres
Publié: (2026)
Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
par: Yang, Zhichao, et autres
Publié: (2026)
par: Yang, Zhichao, et autres
Publié: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
par: Ye, Zhiling, et autres
Publié: (2025)
par: Ye, Zhiling, et autres
Publié: (2025)
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
par: Zhou, Yang, et autres
Publié: (2025)
par: Zhou, Yang, et autres
Publié: (2025)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
par: Huang, Tzu-Heng, et autres
Publié: (2026)
par: Huang, Tzu-Heng, et autres
Publié: (2026)
Sparsely-Supervised Data Assimilation via Physics-Informed Schrödinger Bridge
par: Bu, Dohyun, et autres
Publié: (2026)
par: Bu, Dohyun, et autres
Publié: (2026)
Recursive Agent Optimization
par: Gandhi, Apurva, et autres
Publié: (2026)
par: Gandhi, Apurva, et autres
Publié: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
par: Gunjal, Anisha, et autres
Publié: (2025)
par: Gunjal, Anisha, et autres
Publié: (2025)
Pref-GUIDE: Continual Policy Learning from Real-Time Human Feedback via Preference-Based Learning
par: Ji, Zhengran, et autres
Publié: (2025)
par: Ji, Zhengran, et autres
Publié: (2025)
Sparse Threats, Focused Defense: Criticality-Aware Robust Reinforcement Learning for Safe Autonomous Driving
par: Wei, Qi, et autres
Publié: (2026)
par: Wei, Qi, et autres
Publié: (2026)
SparseDM: Toward Sparse Efficient Diffusion Models
par: Wang, Kafeng, et autres
Publié: (2024)
par: Wang, Kafeng, et autres
Publié: (2024)
Towards Self-Supervised Foundation Models for Critical Care Time Series
par: Jagd, Katja Naasunnguaq, et autres
Publié: (2025)
par: Jagd, Katja Naasunnguaq, et autres
Publié: (2025)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
par: Shen, William F., et autres
Publié: (2026)
par: Shen, William F., et autres
Publié: (2026)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
par: Li, Xingyao, et autres
Publié: (2026)
par: Li, Xingyao, et autres
Publié: (2026)
Locality Sensitive Sparse Encoding for Learning World Models Online
par: Liu, Zichen, et autres
Publié: (2024)
par: Liu, Zichen, et autres
Publié: (2024)
Causally-informed Deep Learning towards Explainable and Generalizable Outcomes Prediction in Critical Care
par: Cheng, Yuxiao, et autres
Publié: (2025)
par: Cheng, Yuxiao, et autres
Publié: (2025)
Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
par: Liu, Genglin, et autres
Publié: (2023)
par: Liu, Genglin, et autres
Publié: (2023)
Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
par: Sholehrasa, Hossein, et autres
Publié: (2025)
par: Sholehrasa, Hossein, et autres
Publié: (2025)
Learning to Detect Critical Nodes in Sparse Graphs via Feature Importance Awareness
par: Tan, Xuwei, et autres
Publié: (2021)
par: Tan, Xuwei, et autres
Publié: (2021)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
par: Hossain, Md Arafat, et autres
Publié: (2025)
par: Hossain, Md Arafat, et autres
Publié: (2025)
Documents similaires
-
Future-as-Label: Scalable Supervision from Real-World Outcomes
par: Turtel, Benjamin, et autres
Publié: (2026) -
Gym-Anything: Turn any Software into an Agent Environment
par: Aggarwal, Pranjal, et autres
Publié: (2026) -
Training Proactive and Personalized LLM Agents
par: Sun, Weiwei, et autres
Publié: (2025) -
CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets
par: Yuan, Lifan, et autres
Publié: (2023) -
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
par: Xie, Lipeng, et autres
Publié: (2025)