GRLO: Towards Generalizable Reinforcement Learning in Open-Ended Environments from Zero
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Shangjian, Fu, Yu, Dong, Yue, Shi, Zhouxing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
G-Zero: Self-Play for Open-Ended Generation from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
von: Mitsides, Konstantinos, et al.
Veröffentlicht: (2026)
von: Mitsides, Konstantinos, et al.
Veröffentlicht: (2026)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
von: Lu, Chris, et al.
Veröffentlicht: (2024)
von: Lu, Chris, et al.
Veröffentlicht: (2024)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
von: Chen, Hui, et al.
Veröffentlicht: (2025)
von: Chen, Hui, et al.
Veröffentlicht: (2025)
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
von: Mannekote, Amogh, et al.
Veröffentlicht: (2024)
von: Mannekote, Amogh, et al.
Veröffentlicht: (2024)
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
von: Nagar, Aishik, et al.
Veröffentlicht: (2026)
von: Nagar, Aishik, et al.
Veröffentlicht: (2026)
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
von: Zhu, Hao, et al.
Veröffentlicht: (2025)
von: Zhu, Hao, et al.
Veröffentlicht: (2025)
DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking
von: Hu, Tianyi, et al.
Veröffentlicht: (2026)
von: Hu, Tianyi, et al.
Veröffentlicht: (2026)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
von: Iyer, Vivek, et al.
Veröffentlicht: (2025)
von: Iyer, Vivek, et al.
Veröffentlicht: (2025)
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
von: Sarangi, Sneheel, et al.
Veröffentlicht: (2025)
von: Sarangi, Sneheel, et al.
Veröffentlicht: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
von: Zheng, Yuxiang, et al.
Veröffentlicht: (2025)
von: Zheng, Yuxiang, et al.
Veröffentlicht: (2025)
KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
von: Duan, Zhangqi, et al.
Veröffentlicht: (2026)
von: Duan, Zhangqi, et al.
Veröffentlicht: (2026)
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
Learning from Failures in Multi-Attempt Reinforcement Learning
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
MCU: An Evaluation Framework for Open-Ended Game Agents
von: Zheng, Xinyue, et al.
Veröffentlicht: (2023)
von: Zheng, Xinyue, et al.
Veröffentlicht: (2023)
RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
RewardAnything: Generalizable Principle-Following Reward Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
Autotelic Reinforcement Learning: Exploring Intrinsic Motivations for Skill Acquisition in Open-Ended Environments
von: Srivastava, Prakhar, et al.
Veröffentlicht: (2025)
von: Srivastava, Prakhar, et al.
Veröffentlicht: (2025)
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
von: Tan, Weihao, et al.
Veröffentlicht: (2024)
von: Tan, Weihao, et al.
Veröffentlicht: (2024)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
von: Wu, Yuhao, et al.
Veröffentlicht: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
von: Almohaimeed, Saad, et al.
Veröffentlicht: (2025)
von: Almohaimeed, Saad, et al.
Veröffentlicht: (2025)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025) -
G-Zero: Self-Play for Open-Ended Generation from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2026) -
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
von: Mitsides, Konstantinos, et al.
Veröffentlicht: (2026) -
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
von: Lu, Chris, et al.
Veröffentlicht: (2024) -
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
von: Chen, Hui, et al.
Veröffentlicht: (2025)