Improving Regret Approximation for Unsupervised Dynamic Environment Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mead, Harry, Lacerda, Bruno, Foerster, Jakob, Hawes, Nick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
von: Mead, Harry, et al.
Veröffentlicht: (2025)
von: Mead, Harry, et al.
Veröffentlicht: (2025)
Refining Minimax Regret for Unsupervised Environment Design
von: Beukman, Michael, et al.
Veröffentlicht: (2024)
von: Beukman, Michael, et al.
Veröffentlicht: (2024)
The Complexity Dynamics of Grokking
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
von: DeMoss, Branton, et al.
Veröffentlicht: (2024)
Monte Carlo Tree Search with Boltzmann Exploration
von: Painter, Michael, et al.
Veröffentlicht: (2024)
von: Painter, Michael, et al.
Veröffentlicht: (2024)
DITTO: Offline Imitation Learning with World Models
von: DeMoss, Branton, et al.
Veröffentlicht: (2023)
von: DeMoss, Branton, et al.
Veröffentlicht: (2023)
JaxWildfire: A GPU-Accelerated Wildfire Simulator for Reinforcement Learning
von: Çakır, Ufuk, et al.
Veröffentlicht: (2025)
von: Çakır, Ufuk, et al.
Veröffentlicht: (2025)
An Optimisation Framework for Unsupervised Environment Design
von: Monette, Nathan, et al.
Veröffentlicht: (2025)
von: Monette, Nathan, et al.
Veröffentlicht: (2025)
Tackling GNARLy Problems: Graph Neural Algorithmic Reasoning Reimagined through Reinforcement Learning
von: Schutz, Alex, et al.
Veröffentlicht: (2025)
von: Schutz, Alex, et al.
Veröffentlicht: (2025)
A Finite-State Controller Based Offline Solver for Deterministic POMDPs
von: Schutz, Alex, et al.
Veröffentlicht: (2025)
von: Schutz, Alex, et al.
Veröffentlicht: (2025)
Improved Dynamic Regret for Online Frank-Wolfe
von: Wan, Yuanyu, et al.
Veröffentlicht: (2023)
von: Wan, Yuanyu, et al.
Veröffentlicht: (2023)
TRACED: Transition-aware Regret Approximation with Co-learnability for Environment Design
von: Cho, Geonwoo, et al.
Veröffentlicht: (2025)
von: Cho, Geonwoo, et al.
Veröffentlicht: (2025)
Discovering Minimal Reinforcement Learning Environments
von: Liesen, Jarek, et al.
Veröffentlicht: (2024)
von: Liesen, Jarek, et al.
Veröffentlicht: (2024)
Improved Approximate Regret for Decentralized Online Continuous Submodular Maximization via Reductions
von: Wan, Yuanyu, et al.
Veröffentlicht: (2026)
von: Wan, Yuanyu, et al.
Veröffentlicht: (2026)
Improving Environment Novelty Quantification for Effective Unsupervised Environment Design
von: Teoh, Jayden, et al.
Veröffentlicht: (2025)
von: Teoh, Jayden, et al.
Veröffentlicht: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
von: Rutherford, Alexander, et al.
Veröffentlicht: (2023)
von: Rutherford, Alexander, et al.
Veröffentlicht: (2023)
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
von: Beukman, Michael, et al.
Veröffentlicht: (2026)
von: Beukman, Michael, et al.
Veröffentlicht: (2026)
Dynamic Regret Reduces to Kernelized Static Regret
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2025)
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2025)
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
von: Ruhdorfer, Constantin, et al.
Veröffentlicht: (2025)
von: Ruhdorfer, Constantin, et al.
Veröffentlicht: (2025)
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
von: Levy, Orin, et al.
Veröffentlicht: (2025)
von: Levy, Orin, et al.
Veröffentlicht: (2025)
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
von: Matthews, Michael, et al.
Veröffentlicht: (2024)
von: Matthews, Michael, et al.
Veröffentlicht: (2024)
JaxUED: A simple and useable UED library in Jax
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Improved Regret of Linear Ensemble Sampling
von: Lee, Harin, et al.
Veröffentlicht: (2024)
von: Lee, Harin, et al.
Veröffentlicht: (2024)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
Follow The Approximate Sparse Leader for No-Regret Online Sparse Linear Approximation
von: Mukhopadhyay, Samrat, et al.
Veröffentlicht: (2025)
von: Mukhopadhyay, Samrat, et al.
Veröffentlicht: (2025)
Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability
von: Zhang, Yu-Jie, et al.
Veröffentlicht: (2025)
von: Zhang, Yu-Jie, et al.
Veröffentlicht: (2025)
Improved Regret Bounds for Bandits with Expert Advice
von: Cesa-Bianchi, Nicolò, et al.
Veröffentlicht: (2024)
von: Cesa-Bianchi, Nicolò, et al.
Veröffentlicht: (2024)
Logarithmic Regret of Exploration in Average Reward Markov Decision Processes
von: Boone, Victor, et al.
Veröffentlicht: (2025)
von: Boone, Victor, et al.
Veröffentlicht: (2025)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive Advantages
von: Jesson, Andrew, et al.
Veröffentlicht: (2023)
von: Jesson, Andrew, et al.
Veröffentlicht: (2023)
Invariance-Based Dynamic Regret Minimization
von: Lazzaretto, Margherita, et al.
Veröffentlicht: (2026)
von: Lazzaretto, Margherita, et al.
Veröffentlicht: (2026)
Improved Regret for Bandit Convex Optimization with Delayed Feedback
von: Wan, Yuanyu, et al.
Veröffentlicht: (2024)
von: Wan, Yuanyu, et al.
Veröffentlicht: (2024)
On Improved Regret Bounds In Bayesian Optimization with Gaussian Noise
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
Quantile-Coupled Flow Matching for Distributional Reinforcement Learning
von: Groom, Michael, et al.
Veröffentlicht: (2026)
von: Groom, Michael, et al.
Veröffentlicht: (2026)
Noisy Zero-Shot Coordination: Breaking The Common Knowledge Assumption In Zero-Shot Coordination Games
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
Solving Robust MDPs through No-Regret Dynamics
von: Guha, Etash Kumar
Veröffentlicht: (2023)
von: Guha, Etash Kumar
Veröffentlicht: (2023)
The Rank-Reduced Kalman Filter: Approximate Dynamical-Low-Rank Filtering In High Dimensions
von: Schmidt, Jonathan, et al.
Veröffentlicht: (2023)
von: Schmidt, Jonathan, et al.
Veröffentlicht: (2023)
Active Context Selection Improves Simple Regret in Contextual Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024) -
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
von: Mead, Harry, et al.
Veröffentlicht: (2025) -
Refining Minimax Regret for Unsupervised Environment Design
von: Beukman, Michael, et al.
Veröffentlicht: (2024) -
The Complexity Dynamics of Grokking
von: DeMoss, Branton, et al.
Veröffentlicht: (2024) -
Monte Carlo Tree Search with Boltzmann Exploration
von: Painter, Michael, et al.
Veröffentlicht: (2024)