RLVR-World: Training World Models with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jialong, Yin, Shaofeng, Feng, Ningya, Long, Mingsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
CompilerDream: Learning a Compiler World Model for General Code Optimization
by: Deng, Chaoyi, et al.
Published: (2024)
by: Deng, Chaoyi, et al.
Published: (2024)
HarmonyDream: Task Harmonization Inside World Models
by: Ma, Haoyu, et al.
Published: (2023)
by: Ma, Haoyu, et al.
Published: (2023)
Trajectory World Models for Heterogeneous Environments
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)
by: Bredis, George, et al.
Published: (2025)
Model-Free Robust Reinforcement Learning with Sample Complexity Analysis
by: Wang, Yudan, et al.
Published: (2024)
by: Wang, Yudan, et al.
Published: (2024)
Spurious Rewards: Rethinking Training Signals in RLVR
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized Worlds
by: Wang, Fan, et al.
Published: (2025)
by: Wang, Fan, et al.
Published: (2025)
In-Context Reinforcement Learning via Communicative World Models
by: Martinez-Lopez, Fernando, et al.
Published: (2025)
by: Martinez-Lopez, Fernando, et al.
Published: (2025)
Continual Reinforcement Learning by Planning with Online World Models
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
Object-Centric World Models for Causality-Aware Reinforcement Learning
by: Nishimoto, Yosuke, et al.
Published: (2025)
by: Nishimoto, Yosuke, et al.
Published: (2025)
Reinforcement Learning from Delayed Observations via World Models
by: Karamzade, Armin, et al.
Published: (2024)
by: Karamzade, Armin, et al.
Published: (2024)
Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models
by: Wu, Hao, et al.
Published: (2025)
by: Wu, Hao, et al.
Published: (2025)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)
by: Lu, Han, et al.
Published: (2025)
Policy and World Modeling Co-Training for Language Agents
by: Lu, Ning, et al.
Published: (2026)
by: Lu, Ning, et al.
Published: (2026)
World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents
by: Fonseca, Yesid, et al.
Published: (2025)
by: Fonseca, Yesid, et al.
Published: (2025)
Novelty Detection in Reinforcement Learning with World Models
by: Zollicoffer, Geigh, et al.
Published: (2023)
by: Zollicoffer, Geigh, et al.
Published: (2023)
From Observations to Events: Event-Aware World Model for Reinforcement Learning
by: Peng, Zhao-Han, et al.
Published: (2026)
by: Peng, Zhao-Han, et al.
Published: (2026)
Vid2World: Crafting Video Diffusion Models to Interactive World Models
by: Huang, Siqiao, et al.
Published: (2025)
by: Huang, Siqiao, et al.
Published: (2025)
Open-World Test-Time Training: Self-Training with Contrast Learning
by: Su, Houcheng, et al.
Published: (2024)
by: Su, Houcheng, et al.
Published: (2024)
Safe Reinforcement Learning for Real-World Engine Control
by: Bedei, Julian, et al.
Published: (2025)
by: Bedei, Julian, et al.
Published: (2025)
Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning
by: Chen, Jiayu, et al.
Published: (2025)
by: Chen, Jiayu, et al.
Published: (2025)
FutureWorld: A Live Reinforcement Learning Environment for Predictive Agents with Real-World Outcome Rewards
by: Han, Zhixin, et al.
Published: (2026)
by: Han, Zhixin, et al.
Published: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR
by: Ingle, Yash, et al.
Published: (2026)
by: Ingle, Yash, et al.
Published: (2026)
Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs
by: Khosravi, Hamed, et al.
Published: (2026)
by: Khosravi, Hamed, et al.
Published: (2026)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
by: Ding, Zihan, et al.
Published: (2024)
by: Ding, Zihan, et al.
Published: (2024)
The Path Not Taken: RLVR Provably Learns Off the Principals
by: Zhu, Hanqing, et al.
Published: (2025)
by: Zhu, Hanqing, et al.
Published: (2025)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
by: Mitsuhashi, Ryo, et al.
Published: (2026)
by: Mitsuhashi, Ryo, et al.
Published: (2026)
Training Agents Inside of Scalable World Models
by: Hafner, Danijar, et al.
Published: (2025)
by: Hafner, Danijar, et al.
Published: (2025)
Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning
by: Qian, Fuyuan, et al.
Published: (2026)
by: Qian, Fuyuan, et al.
Published: (2026)
Real-World Reinforcement Learning of Active Perception Behaviors
by: Hu, Edward S., et al.
Published: (2025)
by: Hu, Edward S., et al.
Published: (2025)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
by: Qiu, Zhongxi, et al.
Published: (2025)
by: Qiu, Zhongxi, et al.
Published: (2025)
DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions
by: Li, Zongyue, et al.
Published: (2025)
by: Li, Zongyue, et al.
Published: (2025)
Real-World Offline Reinforcement Learning from Vision Language Model Feedback
by: Venkataraman, Sreyas, et al.
Published: (2024)
by: Venkataraman, Sreyas, et al.
Published: (2024)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
Dreaming of Many Worlds: Learning Contextual World Models Aids Zero-Shot Generalization
by: Prasanna, Sai, et al.
Published: (2024)
by: Prasanna, Sai, et al.
Published: (2024)
Better World Models Can Lead to Better Post-Training Performance
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
by: Wang, Zhaoyang, et al.
Published: (2026)
by: Wang, Zhaoyang, et al.
Published: (2026)
Optimistic World Models: Efficient Exploration in Model-Based Deep Reinforcement Learning
by: Mete, Akshay, et al.
Published: (2026)
by: Mete, Akshay, et al.
Published: (2026)
Similar Items
-
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024) -
CompilerDream: Learning a Compiler World Model for General Code Optimization
by: Deng, Chaoyi, et al.
Published: (2024) -
HarmonyDream: Task Harmonization Inside World Models
by: Ma, Haoyu, et al.
Published: (2023) -
Trajectory World Models for Heterogeneous Environments
by: Yin, Shaofeng, et al.
Published: (2025) -
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)