Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
Fuente:
arXiv
Saved in:
| Main Authors: | Lan, Guangchen, Xiong, Lian, Zhou, Xin, Cui, Hejie, Zhang, Yuwei, Li, Mao, Shi, Zhenyu, Fetahu, Besnik, Li, Lihong, Li, Xian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
by: Lan, Guangchen
Published: (2026)
by: Lan, Guangchen
Published: (2026)
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024)
by: Hashemi, Helia, et al.
Published: (2024)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
by: Zhou, Yangyang, et al.
Published: (2026)
by: Zhou, Yangyang, et al.
Published: (2026)
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
by: Galvan-Sosa, Diana, et al.
Published: (2025)
by: Galvan-Sosa, Diana, et al.
Published: (2025)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
by: Yuan, Xin, et al.
Published: (2025)
by: Yuan, Xin, et al.
Published: (2025)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
by: Osborne, Philip, et al.
Published: (2025)
by: Osborne, Philip, et al.
Published: (2025)
Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
by: Figueiredo, Vanessa
Published: (2025)
by: Figueiredo, Vanessa
Published: (2025)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
by: Wei, Xingyuan, et al.
Published: (2024)
by: Wei, Xingyuan, et al.
Published: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
by: Steele, Brady, et al.
Published: (2026)
by: Steele, Brady, et al.
Published: (2026)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
by: Aharon, Eliya Naomi, et al.
Published: (2026)
by: Aharon, Eliya Naomi, et al.
Published: (2026)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
by: Sun, Mingrui, et al.
Published: (2026)
by: Sun, Mingrui, et al.
Published: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025)
by: He, Kaifeng, et al.
Published: (2025)
SemanticAgent: A Semantics-Aware Framework for Text-to-SQL Data Synthesis
by: Gao, Qiang, et al.
Published: (2026)
by: Gao, Qiang, et al.
Published: (2026)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023)
by: Gutkin, Alexander, et al.
Published: (2023)
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
by: Chang, Sidi, et al.
Published: (2026)
by: Chang, Sidi, et al.
Published: (2026)
LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning
by: Luijkx, Jelle, et al.
Published: (2025)
by: Luijkx, Jelle, et al.
Published: (2025)
A Pluggable Common Sense-Enhanced Framework for Knowledge Graph Completion
by: Niu, Guanglin, et al.
Published: (2024)
by: Niu, Guanglin, et al.
Published: (2024)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
by: Marmoret, Axel, et al.
Published: (2025)
by: Marmoret, Axel, et al.
Published: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
by: Chen, Yuangong, et al.
Published: (2026)
by: Chen, Yuangong, et al.
Published: (2026)
RefiningGPT: Specialized language Models for Automated Refinery Unit-level Process Diagram Synthesis
by: Liu, Dongxiao, et al.
Published: (2026)
by: Liu, Dongxiao, et al.
Published: (2026)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
by: Chua, Jaymari, et al.
Published: (2025)
by: Chua, Jaymari, et al.
Published: (2025)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
by: Li, Zilin, et al.
Published: (2026)
by: Li, Zilin, et al.
Published: (2026)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
Tensor Generalized Approximate Message Passing
by: Li, Yinchuan, et al.
Published: (2025)
by: Li, Yinchuan, et al.
Published: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
by: Adelson, Trevor, et al.
Published: (2026)
by: Adelson, Trevor, et al.
Published: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
by: Wu, Shuai, et al.
Published: (2026)
by: Wu, Shuai, et al.
Published: (2026)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
by: Amini, Ali
Published: (2025)
by: Amini, Ali
Published: (2025)
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
by: von Cossel, Oskar
Published: (2026)
by: von Cossel, Oskar
Published: (2026)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
by: Fang, Qitong, et al.
Published: (2026)
by: Fang, Qitong, et al.
Published: (2026)
Similar Items
-
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
by: Lan, Guangchen
Published: (2026) -
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
by: Lan, Guangchen, et al.
Published: (2025) -
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024) -
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
by: Lan, Guangchen, et al.
Published: (2025) -
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
by: Zhou, Yangyang, et al.
Published: (2026)