Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Cohen, Lior, Nabati, Ofir, Wang, Kaixin, Kumar, Navdeep, Mannor, Shie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Representation-Driven Reinforcement Learning
di: Nabati, Ofir, et al.
Pubblicazione: (2023)
di: Nabati, Ofir, et al.
Pubblicazione: (2023)
Improving Token-Based World Models with Parallel Observation Prediction
di: Cohen, Lior, et al.
Pubblicazione: (2024)
di: Cohen, Lior, et al.
Pubblicazione: (2024)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
di: Cohen, Lior, et al.
Pubblicazione: (2025)
di: Cohen, Lior, et al.
Pubblicazione: (2025)
Spectral Bellman Method: Unifying Representation and Exploration in RL
di: Nabati, Ofir, et al.
Pubblicazione: (2025)
di: Nabati, Ofir, et al.
Pubblicazione: (2025)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
di: Wang, Kaixin, et al.
Pubblicazione: (2023)
di: Wang, Kaixin, et al.
Pubblicazione: (2023)
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
di: Kumar, Navdeep, et al.
Pubblicazione: (2026)
di: Kumar, Navdeep, et al.
Pubblicazione: (2026)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
di: Koren, Uri, et al.
Pubblicazione: (2025)
di: Koren, Uri, et al.
Pubblicazione: (2025)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
di: Kumar, Navdeep, et al.
Pubblicazione: (2024)
di: Kumar, Navdeep, et al.
Pubblicazione: (2024)
On the Convergence of Single-Timescale Actor-Critic
di: Kumar, Navdeep, et al.
Pubblicazione: (2024)
di: Kumar, Navdeep, et al.
Pubblicazione: (2024)
Efficient Fairness-Performance Pareto Front Computation
di: Kozdoba, Mark, et al.
Pubblicazione: (2024)
di: Kozdoba, Mark, et al.
Pubblicazione: (2024)
MinMaxMin $Q$-learning
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
Conservative DDPG -- Pessimistic RL without Ensemble
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
di: Gadot, Uri, et al.
Pubblicazione: (2023)
di: Gadot, Uri, et al.
Pubblicazione: (2023)
Representative Action Selection for Large Action Space: From Bandits to MDPs
di: Zhou, Quan, et al.
Pubblicazione: (2025)
di: Zhou, Quan, et al.
Pubblicazione: (2025)
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
di: Perets, Binyamin, et al.
Pubblicazione: (2026)
di: Perets, Binyamin, et al.
Pubblicazione: (2026)
Sobolev Space Regularised Pre Density Models
di: Kozdoba, Mark, et al.
Pubblicazione: (2023)
di: Kozdoba, Mark, et al.
Pubblicazione: (2023)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
di: Valensi, David, et al.
Pubblicazione: (2024)
di: Valensi, David, et al.
Pubblicazione: (2024)
The Value of Mechanistic Priors in Sequential Decision Making
di: Shufaro, Itai, et al.
Pubblicazione: (2026)
di: Shufaro, Itai, et al.
Pubblicazione: (2026)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
di: Kumar, Navdeep, et al.
Pubblicazione: (2025)
di: Kumar, Navdeep, et al.
Pubblicazione: (2025)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
di: Du, Yihan, et al.
Pubblicazione: (2024)
di: Du, Yihan, et al.
Pubblicazione: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
di: Kwon, Jeongyeol, et al.
Pubblicazione: (2024)
di: Kwon, Jeongyeol, et al.
Pubblicazione: (2024)
Policy Gradient with Tree Expansion
di: Dalal, Gal, et al.
Pubblicazione: (2023)
di: Dalal, Gal, et al.
Pubblicazione: (2023)
Representative Action Selection for Large Action Space Bandit Families
di: Zhou, Quan, et al.
Pubblicazione: (2025)
di: Zhou, Quan, et al.
Pubblicazione: (2025)
Task Tokens: A Flexible Approach to Adapting Behavior Foundation Models
di: Vainshtein, Ron, et al.
Pubblicazione: (2025)
di: Vainshtein, Ron, et al.
Pubblicazione: (2025)
On Bits and Bandits: Quantifying the Regret-Information Trade-off
di: Shufaro, Itai, et al.
Pubblicazione: (2024)
di: Shufaro, Itai, et al.
Pubblicazione: (2024)
A Classification View on Meta Learning Bandits
di: Mutti, Mirco, et al.
Pubblicazione: (2025)
di: Mutti, Mirco, et al.
Pubblicazione: (2025)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
di: Yoo, Seungwoo, et al.
Pubblicazione: (2026)
di: Yoo, Seungwoo, et al.
Pubblicazione: (2026)
ImagineBench: Evaluating Reinforcement Learning with Large Language Model Rollouts
di: Pang, Jing-Cheng, et al.
Pubblicazione: (2025)
di: Pang, Jing-Cheng, et al.
Pubblicazione: (2025)
Reinforcement Learning with Segment Feedback
di: Du, Yihan, et al.
Pubblicazione: (2025)
di: Du, Yihan, et al.
Pubblicazione: (2025)
SQT -- std $Q$-target
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
DyDiff: Long-Horizon Rollout via Dynamics Diffusion for Offline Reinforcement Learning
di: Zhao, Hanye, et al.
Pubblicazione: (2024)
di: Zhao, Hanye, et al.
Pubblicazione: (2024)
VLM-Guided Experience Replay
di: Sharony, Elad, et al.
Pubblicazione: (2026)
di: Sharony, Elad, et al.
Pubblicazione: (2026)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
di: Zhai, Yuanzhao, et al.
Pubblicazione: (2024)
di: Zhai, Yuanzhao, et al.
Pubblicazione: (2024)
State Entropy Regularization for Robust Reinforcement Learning
di: Ashlag, Yonatan, et al.
Pubblicazione: (2025)
di: Ashlag, Yonatan, et al.
Pubblicazione: (2025)
Adversarial Bandit over Bandits: Hierarchical Bandits for Online Configuration Management
di: Avin, Chen, et al.
Pubblicazione: (2025)
di: Avin, Chen, et al.
Pubblicazione: (2025)
RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
di: Gadot, Uri, et al.
Pubblicazione: (2025)
di: Gadot, Uri, et al.
Pubblicazione: (2025)
SGNO: Spectral Generator Neural Operators for Stable Long Horizon PDE Rollouts
di: Li, Jiayi, et al.
Pubblicazione: (2026)
di: Li, Jiayi, et al.
Pubblicazione: (2026)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
di: Ding, Zihan, et al.
Pubblicazione: (2024)
di: Ding, Zihan, et al.
Pubblicazione: (2024)
Learning Multiple Initial Solutions to Optimization Problems
di: Sharony, Elad, et al.
Pubblicazione: (2024)
di: Sharony, Elad, et al.
Pubblicazione: (2024)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
di: Mundada, Gagan, et al.
Pubblicazione: (2026)
di: Mundada, Gagan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Representation-Driven Reinforcement Learning
di: Nabati, Ofir, et al.
Pubblicazione: (2023) -
Improving Token-Based World Models with Parallel Observation Prediction
di: Cohen, Lior, et al.
Pubblicazione: (2024) -
Simulus: Combining Improvements in Sample-Efficient World Model Agents
di: Cohen, Lior, et al.
Pubblicazione: (2025) -
Spectral Bellman Method: Unifying Representation and Exploration in RL
di: Nabati, Ofir, et al.
Pubblicazione: (2025) -
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
di: Wang, Kaixin, et al.
Pubblicazione: (2023)