Artificial Replay: A Meta-Algorithm for Harnessing Historical Data in Bandits
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Banerjee, Siddhartha, Sinclair, Sean R., Tambe, Milind, Xu, Lily, Yu, Christina Lee |
|---|---|
| Format: | Preprint |
| Publié: |
2022
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits
par: Liang, Biyonka, et autres
Publié: (2024)
par: Liang, Biyonka, et autres
Publié: (2024)
Adaptive Discretization in Online Reinforcement Learning
par: Sinclair, Sean R., et autres
Publié: (2021)
par: Sinclair, Sean R., et autres
Publié: (2021)
Dual-Mandate Patrols: Multi-Armed Bandits for Green Security
par: Xu, Lily, et autres
Publié: (2020)
par: Xu, Lily, et autres
Publié: (2020)
Combining Diverse Information for Coordinated Action: Stochastic Bandit Algorithms for Heterogeneous Agents
par: Gordon, Lucia, et autres
Publié: (2024)
par: Gordon, Lucia, et autres
Publié: (2024)
The Bandit Whisperer: Communication Learning for Restless Bandits
par: Zhao, Yunfan, et autres
Publié: (2024)
par: Zhao, Yunfan, et autres
Publié: (2024)
Reinforcement learning with combinatorial actions for coupled restless bandits
par: Xu, Lily, et autres
Publié: (2025)
par: Xu, Lily, et autres
Publié: (2025)
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
par: Verma, Shresth, et autres
Publié: (2024)
par: Verma, Shresth, et autres
Publié: (2024)
Online Fair Allocation of Perishable Resources
par: Banerjee, Siddhartha, et autres
Publié: (2024)
par: Banerjee, Siddhartha, et autres
Publié: (2024)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
par: Jain, Gauri, et autres
Publié: (2024)
par: Jain, Gauri, et autres
Publié: (2024)
Bayesian Collaborative Bandits with Thompson Sampling for Improved Outreach in Maternal Health Program
par: Dasgupta, Arpan, et autres
Publié: (2024)
par: Dasgupta, Arpan, et autres
Publié: (2024)
Decisions and Deployment: The Five-Year SAHELI Project (2020-2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health
par: Verma, Shresth, et autres
Publié: (2026)
par: Verma, Shresth, et autres
Publié: (2026)
Analyzing Cost-Sensitive Surrogate Losses via $\mathcal{H}$-calibration
par: Shah, Sanket, et autres
Publié: (2025)
par: Shah, Sanket, et autres
Publié: (2025)
A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health
par: Behari, Nikhil, et autres
Publié: (2024)
par: Behari, Nikhil, et autres
Publié: (2024)
Fairness for Workers Who Pull the Arms: An Index Based Policy for Allocation of Restless Bandit Tasks
par: Biswas, Arpita, et autres
Publié: (2023)
par: Biswas, Arpita, et autres
Publié: (2023)
Finite-Horizon Single-Pull Restless Bandits: An Efficient Index Policy For Scarce Resource Allocation
par: Xiong, Guojun, et autres
Publié: (2025)
par: Xiong, Guojun, et autres
Publié: (2025)
The SMART approach to instance-optimal online learning
par: Banerjee, Siddhartha, et autres
Publié: (2024)
par: Banerjee, Siddhartha, et autres
Publié: (2024)
Towards a Pretrained Model for Restless Bandits via Multi-arm Generalization
par: Zhao, Yunfan, et autres
Publié: (2023)
par: Zhao, Yunfan, et autres
Publié: (2023)
Generative AI Against Poaching: Latent Composite Flow Matching for Wildlife Conservation
par: Kong, Lingkai, et autres
Publié: (2025)
par: Kong, Lingkai, et autres
Publié: (2025)
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
par: Kong, Lingkai, et autres
Publié: (2025)
par: Kong, Lingkai, et autres
Publié: (2025)
The Data-Driven Censored Newsvendor Problem
par: Hssaine, Chamsi, et autres
Publié: (2024)
par: Hssaine, Chamsi, et autres
Publié: (2024)
Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions
par: Kong, Lingkai, et autres
Publié: (2026)
par: Kong, Lingkai, et autres
Publié: (2026)
Improving Health Information Access in the World's Largest Maternal Mobile Health Program via Bandit Algorithms
par: Lalan, Arshika, et autres
Publié: (2024)
par: Lalan, Arshika, et autres
Publié: (2024)
Lightweight Robust Direct Preference Optimization
par: Kim, Cheol Woo, et autres
Publié: (2025)
par: Kim, Cheol Woo, et autres
Publié: (2025)
Preference Robustness for DPO with Applications to Public Health
par: Kim, Cheol Woo, et autres
Publié: (2025)
par: Kim, Cheol Woo, et autres
Publié: (2025)
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
par: Parthasarathy, Ambreesh, et autres
Publié: (2025)
par: Parthasarathy, Ambreesh, et autres
Publié: (2025)
Leaving the Nest: Going Beyond Local Loss Functions for Predict-Then-Optimize
par: Shah, Sanket, et autres
Publié: (2023)
par: Shah, Sanket, et autres
Publié: (2023)
Efficient Public Health Intervention Planning Using Decomposition-Based Decision-Focused Learning
par: Shah, Sanket, et autres
Publié: (2024)
par: Shah, Sanket, et autres
Publié: (2024)
A Reduction Algorithm for Markovian Contextual Linear Bandits
par: Buyukkalayci, Kaan, et autres
Publié: (2026)
par: Buyukkalayci, Kaan, et autres
Publié: (2026)
SelfReplay: Adapting Self-Supervised Sensory Models via Adaptive Meta-Task Replay
par: Yoon, Hyungjun, et autres
Publié: (2024)
par: Yoon, Hyungjun, et autres
Publié: (2024)
Contrasting local and global modeling with machine learning and satellite data: A case study estimating tree canopy height in African savannas
par: Rolf, Esther, et autres
Publié: (2024)
par: Rolf, Esther, et autres
Publié: (2024)
On Diffusion Models for Multi-Agent Partial Observability: Shared Attractors, Error Bounds, and Composite Flow
par: Wang, Tonghan, et autres
Publié: (2024)
par: Wang, Tonghan, et autres
Publié: (2024)
Improving the Prediction of Individual Engagement in Recommendations Using Cognitive Models
par: Seow, Roderick, et autres
Publié: (2024)
par: Seow, Roderick, et autres
Publié: (2024)
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
par: Tec, Mauricio, et autres
Publié: (2025)
par: Tec, Mauricio, et autres
Publié: (2025)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
par: Zhang, Zhongjun, et autres
Publié: (2026)
par: Zhang, Zhongjun, et autres
Publié: (2026)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
par: Wang, Haichuan, et autres
Publié: (2026)
par: Wang, Haichuan, et autres
Publié: (2026)
Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
par: Ai, Rui, et autres
Publié: (2025)
par: Ai, Rui, et autres
Publié: (2025)
Optimal and Practical Batched Linear Bandit Algorithm
par: Yu, Sanghoon, et autres
Publié: (2025)
par: Yu, Sanghoon, et autres
Publié: (2025)
Reinforcement Learning in MDPs with Information-Ordered Policies
par: Zhang, Zhongjun, et autres
Publié: (2025)
par: Zhang, Zhongjun, et autres
Publié: (2025)
What is the Right Notion of Distance between Predict-then-Optimize Tasks?
par: Rodriguez-Diaz, Paula, et autres
Publié: (2024)
par: Rodriguez-Diaz, Paula, et autres
Publié: (2024)
A Classification View on Meta Learning Bandits
par: Mutti, Mirco, et autres
Publié: (2025)
par: Mutti, Mirco, et autres
Publié: (2025)
Documents similaires
-
Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits
par: Liang, Biyonka, et autres
Publié: (2024) -
Adaptive Discretization in Online Reinforcement Learning
par: Sinclair, Sean R., et autres
Publié: (2021) -
Dual-Mandate Patrols: Multi-Armed Bandits for Green Security
par: Xu, Lily, et autres
Publié: (2020) -
Combining Diverse Information for Coordinated Action: Stochastic Bandit Algorithms for Heterogeneous Agents
par: Gordon, Lucia, et autres
Publié: (2024) -
The Bandit Whisperer: Communication Learning for Restless Bandits
par: Zhao, Yunfan, et autres
Publié: (2024)