Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Mukherjee, Subhojyoti, Hanna, Josiah P., Xie, Qiaomin, Nowak, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
Stable Offline Value Function Learning with Bisimulation-based Representations
by: Pavse, Brahma S., et al.
Published: (2024)
by: Pavse, Brahma S., et al.
Published: (2024)
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
by: Pavse, Brahma S., et al.
Published: (2023)
by: Pavse, Brahma S., et al.
Published: (2023)
Efficient and Interpretable Bandit Algorithms
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
An Empirical Study on the Power of Future Prediction in Partially Observable Environments
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
In-Context Curiosity: Distilling Exploration for Decision-Pretrained Transformers on Bandit Tasks
by: Yang, Huitao, et al.
Published: (2025)
by: Yang, Huitao, et al.
Published: (2025)
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2023)
by: Hong, Yige, et al.
Published: (2023)
Achieving Exponential Asymptotic Optimality in Average-Reward Restless Bandits without Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation
by: Yan, Hao, et al.
Published: (2025)
by: Yan, Hao, et al.
Published: (2025)
Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies
by: Corrado, Nicholas E., et al.
Published: (2025)
by: Corrado, Nicholas E., et al.
Published: (2025)
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling
by: Corrado, Nicholas E., et al.
Published: (2026)
by: Corrado, Nicholas E., et al.
Published: (2026)
Prompt Tuning Decision Transformers with Structured and Scalable Bandits
by: Rietz, Finn, et al.
Published: (2025)
by: Rietz, Finn, et al.
Published: (2025)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
by: Corrado, Nicholas E., et al.
Published: (2023)
by: Corrado, Nicholas E., et al.
Published: (2023)
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates
by: Corrado, Nicholas E., et al.
Published: (2023)
by: Corrado, Nicholas E., et al.
Published: (2023)
Wasserstein-p Central Limit Theorem Rates: From Local Dependence to Markov Chains
by: Zhang, Yixuan, et al.
Published: (2026)
by: Zhang, Yixuan, et al.
Published: (2026)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Contextual Online Pricing with (Biased) Offline Data
by: Zhang, Yixuan, et al.
Published: (2025)
by: Zhang, Yixuan, et al.
Published: (2025)
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
by: Zhang, Yixuan, et al.
Published: (2026)
by: Zhang, Yixuan, et al.
Published: (2026)
Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards
by: Ji, Wenlong, et al.
Published: (2025)
by: Ji, Wenlong, et al.
Published: (2025)
Monotonic Transformation Invariant Multi-task Learning
by: Murthy, Surya, et al.
Published: (2025)
by: Murthy, Surya, et al.
Published: (2025)
Agentic Planning with Reasoning for Image Styling via Offline RL
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
BanditQ: Fair Bandits with Guaranteed Rewards
by: Sinha, Abhishek
Published: (2023)
by: Sinha, Abhishek
Published: (2023)
Transfer in Sequential Multi-armed Bandits via Reward Samples
by: R, Rahul N, et al.
Published: (2024)
by: R, Rahul N, et al.
Published: (2024)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
by: Huo, Dongyan, et al.
Published: (2022)
by: Huo, Dongyan, et al.
Published: (2022)
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
by: Mukherjee, Subhojyoti, et al.
Published: (2025)
by: Mukherjee, Subhojyoti, et al.
Published: (2025)
Learning for Bandits under Action Erasures
by: Hanna, Osama, et al.
Published: (2024)
by: Hanna, Osama, et al.
Published: (2024)
Off-Policy Evaluation from Logged Human Feedback
by: Bhargava, Aniruddha, et al.
Published: (2024)
by: Bhargava, Aniruddha, et al.
Published: (2024)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
by: Nguyen, Duy, et al.
Published: (2024)
by: Nguyen, Duy, et al.
Published: (2024)
Latent Diffusion Pretraining for Crystal Property Prediction
by: Mukherjee, Shrimon, et al.
Published: (2026)
by: Mukherjee, Shrimon, et al.
Published: (2026)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
by: Jain, Arushi, et al.
Published: (2024)
by: Jain, Arushi, et al.
Published: (2024)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
by: Qin, Hao, et al.
Published: (2023)
by: Qin, Hao, et al.
Published: (2023)
Global Rewards in Restless Multi-Armed Bandits
by: Raman, Naveen, et al.
Published: (2024)
by: Raman, Naveen, et al.
Published: (2024)
Similar Items
-
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
by: Mukherjee, Subhojyoti, et al.
Published: (2023) -
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
by: Mukherjee, Subhojyoti, et al.
Published: (2024) -
Stable Offline Value Function Learning with Bisimulation-based Representations
by: Pavse, Brahma S., et al.
Published: (2024) -
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
by: Pavse, Brahma S., et al.
Published: (2023) -
Efficient and Interpretable Bandit Algorithms
by: Mukherjee, Subhojyoti, et al.
Published: (2023)