On the Hardness of Reinforcement Learning with Transition Look-Ahead
Fuente:
arXiv
Saved in:
| Main Authors: | Pla, Corentin, Richard, Hugo, Abeille, Marc, Merlis, Nadav, Perchet, Vianney |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Stable Matching with Ties: Approximation Ratios and Learning
by: Lin, Shiyun, et al.
Published: (2024)
by: Lin, Shiyun, et al.
Published: (2024)
Adaptive Bandit Algorithms for Contextual Matching Markets
by: Lin, Shiyun, et al.
Published: (2026)
by: Lin, Shiyun, et al.
Published: (2026)
Improved Algorithms for Contextual Dynamic Pricing
by: Tullii, Matilde, et al.
Published: (2024)
by: Tullii, Matilde, et al.
Published: (2024)
Reinforcement Learning with Lookahead Information
by: Merlis, Nadav
Published: (2024)
by: Merlis, Nadav
Published: (2024)
Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching
by: Merlis, Nadav
Published: (2026)
by: Merlis, Nadav
Published: (2026)
Multi-Armed Bandits with Minimum Aggregated Revenue Constraints
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
Learning to Allocate Resources with Censored Feedback
by: Montanari, Giovanni, et al.
Published: (2026)
by: Montanari, Giovanni, et al.
Published: (2026)
Learning in Prophet Inequalities with Noisy Observations
by: Kim, Jung-hun, et al.
Published: (2026)
by: Kim, Jung-hun, et al.
Published: (2026)
Instance-dependent Stochastic Lipschitz bandit
by: Potfer, Marius, et al.
Published: (2026)
by: Potfer, Marius, et al.
Published: (2026)
On Tradeoffs in Learning-Augmented Algorithms
by: Benomar, Ziyad, et al.
Published: (2025)
by: Benomar, Ziyad, et al.
Published: (2025)
Online Packet Scheduling with Deadlines and Learning
by: Genalti, Gianmarco, et al.
Published: (2026)
by: Genalti, Gianmarco, et al.
Published: (2026)
Comparing Uniform Price and Discriminatory Multi-Unit Auctions through Regret Minimization
by: Potfer, Marius, et al.
Published: (2025)
by: Potfer, Marius, et al.
Published: (2025)
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Non-clairvoyant Scheduling with Partial Predictions
by: Benomar, Ziyad, et al.
Published: (2024)
by: Benomar, Ziyad, et al.
Published: (2024)
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
by: Perrault, Pierre, et al.
Published: (2026)
by: Perrault, Pierre, et al.
Published: (2026)
Improved learning rates in multi-unit uniform price auctions
by: Potfer, Marius, et al.
Published: (2025)
by: Potfer, Marius, et al.
Published: (2025)
Online Linear Regression with Paid Stochastic Features
by: Merlis, Nadav, et al.
Published: (2025)
by: Merlis, Nadav, et al.
Published: (2025)
Distribution-Aware Mean Estimation under User-level Local Differential Privacy
by: Pla, Corentin, et al.
Published: (2024)
by: Pla, Corentin, et al.
Published: (2024)
Look Before Leap: Look-Ahead Planning with Uncertainty in Reinforcement Learning
by: Liu, Yongshuai, et al.
Published: (2025)
by: Liu, Yongshuai, et al.
Published: (2025)
On Bits and Bandits: Quantifying the Regret-Information Trade-off
by: Shufaro, Itai, et al.
Published: (2024)
by: Shufaro, Itai, et al.
Published: (2024)
Variance-sensitive Thompson sampling for generalised linear bandits, revisited
by: Perneczky, Tom, et al.
Published: (2026)
by: Perneczky, Tom, et al.
Published: (2026)
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
by: Fiegel, Côme, et al.
Published: (2026)
by: Fiegel, Côme, et al.
Published: (2026)
Strategic Multi-Armed Bandit Problems Under Debt-Free Reporting
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
When and why randomised exploration works (in linear bandits)
by: Abeille, Marc, et al.
Published: (2025)
by: Abeille, Marc, et al.
Published: (2025)
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
Looking Ahead to Avoid Being Late: Solving Hard-Constrained Traveling Salesman Problem
by: Chen, Jingxiao, et al.
Published: (2024)
by: Chen, Jingxiao, et al.
Published: (2024)
Look-Ahead Reasoning on Learning Platforms
by: Zhu, Haiqing, et al.
Published: (2025)
by: Zhu, Haiqing, et al.
Published: (2025)
Mode Estimation with Partial Feedback
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
by: Jenner, Erik, et al.
Published: (2024)
by: Jenner, Erik, et al.
Published: (2024)
The Equilibrium Response of Atmospheric Machine-Learning Models to Uniform Sea Surface Temperature Warming
by: Zhang, Bosong, et al.
Published: (2025)
by: Zhang, Bosong, et al.
Published: (2025)
Streaming Looking Ahead with Token-level Self-reward
by: Zhang, Hongming, et al.
Published: (2025)
by: Zhang, Hongming, et al.
Published: (2025)
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
DU-Shapley: A Shapley Value Proxy for Efficient Dataset Valuation
by: Garrido-Lucero, Felipe, et al.
Published: (2023)
by: Garrido-Lucero, Felipe, et al.
Published: (2023)
Look-Ahead Screening Rules for the Lasso
by: Larsson, Johan
Published: (2021)
by: Larsson, Johan
Published: (2021)
LAMP: Look-Ahead Mixed-Precision Inference of Large Language Models
by: Budzinskiy, Stanislav, et al.
Published: (2026)
by: Budzinskiy, Stanislav, et al.
Published: (2026)
FigBO: A Generalized Acquisition Function Framework with Look-Ahead Capability for Bayesian Optimization
by: Chen, Hui, et al.
Published: (2025)
by: Chen, Hui, et al.
Published: (2025)
Calibrated Forecasting and Persuasion
by: Jain, Atulya, et al.
Published: (2024)
by: Jain, Atulya, et al.
Published: (2024)
Graph-based Semi-Supervised Learning via Maximum Discrimination
by: Katz, Nadav, et al.
Published: (2026)
by: Katz, Nadav, et al.
Published: (2026)
Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance
by: Benhenda, Mostapha
Published: (2026)
by: Benhenda, Mostapha
Published: (2026)
Similar Items
-
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024) -
Stable Matching with Ties: Approximation Ratios and Learning
by: Lin, Shiyun, et al.
Published: (2024) -
Adaptive Bandit Algorithms for Contextual Matching Markets
by: Lin, Shiyun, et al.
Published: (2026) -
Improved Algorithms for Contextual Dynamic Pricing
by: Tullii, Matilde, et al.
Published: (2024) -
Reinforcement Learning with Lookahead Information
by: Merlis, Nadav
Published: (2024)