Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Xiyan, Liu, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring Continual Learning of Compositional Generalization in NLI
por: Fu, Xiyan, et al.
Publicado: (2024)
por: Fu, Xiyan, et al.
Publicado: (2024)
The Mystery of Compositional Generalization in Graph-based Generative Commonsense Reasoning
por: Fu, Xiyan, et al.
Publicado: (2024)
por: Fu, Xiyan, et al.
Publicado: (2024)
IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
por: Baek, In-Chang, et al.
Publicado: (2025)
por: Baek, In-Chang, et al.
Publicado: (2025)
Compositional Instruction Following with Language Models and Reinforcement Learning
por: Cohen, Vanya, et al.
Publicado: (2025)
por: Cohen, Vanya, et al.
Publicado: (2025)
Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
por: Liu, Chi, et al.
Publicado: (2025)
por: Liu, Chi, et al.
Publicado: (2025)
General Preference Reinforcement Learning
por: Umer, Muhammad, et al.
Publicado: (2026)
por: Umer, Muhammad, et al.
Publicado: (2026)
MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning
por: Wang, Hongjun, et al.
Publicado: (2026)
por: Wang, Hongjun, et al.
Publicado: (2026)
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
por: Zhao, Yang, et al.
Publicado: (2025)
por: Zhao, Yang, et al.
Publicado: (2025)
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
por: Jiang, Guochao, et al.
Publicado: (2026)
por: Jiang, Guochao, et al.
Publicado: (2026)
Complementary Reinforcement Learning
por: Muhtar, Dilxat, et al.
Publicado: (2026)
por: Muhtar, Dilxat, et al.
Publicado: (2026)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
por: Bai, Yang, et al.
Publicado: (2026)
por: Bai, Yang, et al.
Publicado: (2026)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
por: Liu, Wei, et al.
Publicado: (2026)
por: Liu, Wei, et al.
Publicado: (2026)
Reinforcement Learning for Tool-Integrated Interleaved Thinking towards Cross-Domain Generalization
por: Chen, Zhengyu, et al.
Publicado: (2025)
por: Chen, Zhengyu, et al.
Publicado: (2025)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
por: Xu, Wujiang, et al.
Publicado: (2025)
por: Xu, Wujiang, et al.
Publicado: (2025)
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
por: Shandilya, Shivam, et al.
Publicado: (2024)
por: Shandilya, Shivam, et al.
Publicado: (2024)
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning
por: Li, Ran, et al.
Publicado: (2026)
por: Li, Ran, et al.
Publicado: (2026)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
por: Han, Sungjun, et al.
Publicado: (2024)
por: Han, Sungjun, et al.
Publicado: (2024)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
por: He, Shenghua, et al.
Publicado: (2025)
por: He, Shenghua, et al.
Publicado: (2025)
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
por: Ding, Fei, et al.
Publicado: (2026)
por: Ding, Fei, et al.
Publicado: (2026)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
por: Lyu, Chengqi, et al.
Publicado: (2025)
por: Lyu, Chengqi, et al.
Publicado: (2025)
Learning from Failures in Multi-Attempt Reinforcement Learning
por: Chung, Stephen, et al.
Publicado: (2025)
por: Chung, Stephen, et al.
Publicado: (2025)
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
por: Li, Junliang, et al.
Publicado: (2025)
por: Li, Junliang, et al.
Publicado: (2025)
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
por: Ou, Jiefu, et al.
Publicado: (2026)
por: Ou, Jiefu, et al.
Publicado: (2026)
ToolExpander: Extending the Frontiers of Tool-Using Reinforcement Learning to Weak LLMs
por: Chen, Fu, et al.
Publicado: (2025)
por: Chen, Fu, et al.
Publicado: (2025)
Reinforcing General Reasoning without Verifiers
por: Zhou, Xiangxin, et al.
Publicado: (2025)
por: Zhou, Xiangxin, et al.
Publicado: (2025)
Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers
por: Hu, Senkang, et al.
Publicado: (2026)
por: Hu, Senkang, et al.
Publicado: (2026)
Reinforced Attention Learning
por: Li, Bangzheng, et al.
Publicado: (2026)
por: Li, Bangzheng, et al.
Publicado: (2026)
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
por: Wang, Yikai, et al.
Publicado: (2026)
por: Wang, Yikai, et al.
Publicado: (2026)
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
por: Su, Zeli, et al.
Publicado: (2026)
por: Su, Zeli, et al.
Publicado: (2026)
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
por: Su, Zhenpeng, et al.
Publicado: (2025)
por: Su, Zhenpeng, et al.
Publicado: (2025)
On Provable Length and Compositional Generalization
por: Ahuja, Kartik, et al.
Publicado: (2024)
por: Ahuja, Kartik, et al.
Publicado: (2024)
Natural Language Reinforcement Learning
por: Feng, Xidong, et al.
Publicado: (2024)
por: Feng, Xidong, et al.
Publicado: (2024)
LLM-Upgraded Graph Reinforcement Learning for Carbon-Aware Job Scheduling in Smart Manufacturing
por: Yang, Zhiying, et al.
Publicado: (2025)
por: Yang, Zhiying, et al.
Publicado: (2025)
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
por: Yu, Qiying, et al.
Publicado: (2025)
por: Yu, Qiying, et al.
Publicado: (2025)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
por: Zhang, Xin, et al.
Publicado: (2026)
por: Zhang, Xin, et al.
Publicado: (2026)
Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation
por: Min, Do June, et al.
Publicado: (2024)
por: Min, Do June, et al.
Publicado: (2024)
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning
por: Yalcinkaya, Beyazit, et al.
Publicado: (2024)
por: Yalcinkaya, Beyazit, et al.
Publicado: (2024)
How Reliable is Multilingual LLM-as-a-Judge?
por: Fu, Xiyan, et al.
Publicado: (2025)
por: Fu, Xiyan, et al.
Publicado: (2025)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
por: Feng, Weitao, et al.
Publicado: (2025)
por: Feng, Weitao, et al.
Publicado: (2025)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
por: Xia, Yinan, et al.
Publicado: (2026)
por: Xia, Yinan, et al.
Publicado: (2026)
Ejemplares similares
-
Exploring Continual Learning of Compositional Generalization in NLI
por: Fu, Xiyan, et al.
Publicado: (2024) -
The Mystery of Compositional Generalization in Graph-based Generative Commonsense Reasoning
por: Fu, Xiyan, et al.
Publicado: (2024) -
IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
por: Baek, In-Chang, et al.
Publicado: (2025) -
Compositional Instruction Following with Language Models and Reinforcement Learning
por: Cohen, Vanya, et al.
Publicado: (2025) -
Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
por: Liu, Chi, et al.
Publicado: (2025)