On Bits and Bandits: Quantifying the Regret-Information Trade-off
Fuente:
arXiv
Saved in:
| Main Authors: | Shufaro, Itai, Merlis, Nadav, Weinberger, Nir, Mannor, Shie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Value of Mechanistic Priors in Sequential Decision Making
by: Shufaro, Itai, et al.
Published: (2026)
by: Shufaro, Itai, et al.
Published: (2026)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Reinforcement Learning with Lookahead Information
by: Merlis, Nadav
Published: (2024)
by: Merlis, Nadav
Published: (2024)
Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching
by: Merlis, Nadav
Published: (2026)
by: Merlis, Nadav
Published: (2026)
Representative Action Selection for Large Action Space: From Bandits to MDPs
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
Representative Action Selection for Large Action Space Bandit Families
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
A Classification View on Meta Learning Bandits
by: Mutti, Mirco, et al.
Published: (2025)
by: Mutti, Mirco, et al.
Published: (2025)
Adversarial Bandit over Bandits: Hierarchical Bandits for Online Configuration Management
by: Avin, Chen, et al.
Published: (2025)
by: Avin, Chen, et al.
Published: (2025)
MinMaxMin $Q$-learning
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Conservative DDPG -- Pessimistic RL without Ensemble
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Adaptive Bandit Algorithms for Contextual Matching Markets
by: Lin, Shiyun, et al.
Published: (2026)
by: Lin, Shiyun, et al.
Published: (2026)
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
by: Perets, Binyamin, et al.
Published: (2026)
by: Perets, Binyamin, et al.
Published: (2026)
Efficient Fairness-Performance Pareto Front Computation
by: Kozdoba, Mark, et al.
Published: (2024)
by: Kozdoba, Mark, et al.
Published: (2024)
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Representation-Driven Reinforcement Learning
by: Nabati, Ofir, et al.
Published: (2023)
by: Nabati, Ofir, et al.
Published: (2023)
Sobolev Space Regularised Pre Density Models
by: Kozdoba, Mark, et al.
Published: (2023)
by: Kozdoba, Mark, et al.
Published: (2023)
Online Linear Regression with Paid Stochastic Features
by: Merlis, Nadav, et al.
Published: (2025)
by: Merlis, Nadav, et al.
Published: (2025)
Spectral Bellman Method: Unifying Representation and Exploration in RL
by: Nabati, Ofir, et al.
Published: (2025)
by: Nabati, Ofir, et al.
Published: (2025)
Task Tokens: A Flexible Approach to Adapting Behavior Foundation Models
by: Vainshtein, Ron, et al.
Published: (2025)
by: Vainshtein, Ron, et al.
Published: (2025)
Regret Distribution in Stochastic Bandits: Optimal Trade-off between Expectation and Tail Risk
by: Simchi-Levi, David, et al.
Published: (2023)
by: Simchi-Levi, David, et al.
Published: (2023)
Improving Token-Based World Models with Parallel Observation Prediction
by: Cohen, Lior, et al.
Published: (2024)
by: Cohen, Lior, et al.
Published: (2024)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
by: Valensi, David, et al.
Published: (2024)
by: Valensi, David, et al.
Published: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
by: Cohen, Lior, et al.
Published: (2026)
by: Cohen, Lior, et al.
Published: (2026)
Reinforcement Learning with Segment Feedback
by: Du, Yihan, et al.
Published: (2025)
by: Du, Yihan, et al.
Published: (2025)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
by: Wang, Kaixin, et al.
Published: (2023)
by: Wang, Kaixin, et al.
Published: (2023)
SQT -- std $Q$-target
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Stable Matching with Ties: Approximation Ratios and Learning
by: Lin, Shiyun, et al.
Published: (2024)
by: Lin, Shiyun, et al.
Published: (2024)
On the Hardness of Reinforcement Learning with Transition Look-Ahead
by: Pla, Corentin, et al.
Published: (2025)
by: Pla, Corentin, et al.
Published: (2025)
On the Convergence of Single-Timescale Actor-Critic
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
by: Cohen, Lior, et al.
Published: (2025)
by: Cohen, Lior, et al.
Published: (2025)
Information Capacity Regret Bounds for Bandits with Mediator Feedback
by: Eldowa, Khaled, et al.
Published: (2024)
by: Eldowa, Khaled, et al.
Published: (2024)
Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation
by: Cohen, Nadav Z., et al.
Published: (2024)
by: Cohen, Nadav Z., et al.
Published: (2024)
Improved Algorithms for Contextual Dynamic Pricing
by: Tullii, Matilde, et al.
Published: (2024)
by: Tullii, Matilde, et al.
Published: (2024)
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026)
by: Sharony, Elad, et al.
Published: (2026)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
Learning Flock: Enhancing Sets of Particles for Multi~Sub-State Particle Filtering with Neural Augmentation
by: Nuri, Itai, et al.
Published: (2024)
by: Nuri, Itai, et al.
Published: (2024)
Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical Inference
by: Wang, Zichen, et al.
Published: (2025)
by: Wang, Zichen, et al.
Published: (2025)
A representation-learning game for classes of prediction tasks
by: Uzan, Neria, et al.
Published: (2024)
by: Uzan, Neria, et al.
Published: (2024)
Similar Items
-
The Value of Mechanistic Priors in Sequential Decision Making
by: Shufaro, Itai, et al.
Published: (2026) -
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2024) -
Reinforcement Learning with Lookahead Information
by: Merlis, Nadav
Published: (2024) -
Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching
by: Merlis, Nadav
Published: (2026) -
Representative Action Selection for Large Action Space: From Bandits to MDPs
by: Zhou, Quan, et al.
Published: (2025)