Q-learning with Posterior Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Agrawal, Priyank, Agrawal, Shipra, Azati, Azmat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimistic Q-learning for average reward and episodic reinforcement learning
by: Agrawal, Priyank, et al.
Published: (2024)
by: Agrawal, Priyank, et al.
Published: (2024)
Dynamic Pricing and Learning with Long-term Reference Effects
by: Agrawal, Shipra, et al.
Published: (2024)
by: Agrawal, Shipra, et al.
Published: (2024)
A Tractable Online Learning Algorithm for the Multinomial Logit Contextual Bandit
by: Agrawal, Priyank, et al.
Published: (2020)
by: Agrawal, Priyank, et al.
Published: (2020)
Dynamic Pricing and Advertising with Demand Learning
by: Agrawal, Shipra, et al.
Published: (2023)
by: Agrawal, Shipra, et al.
Published: (2023)
Spectral Thompson sampling
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Spectral bandits for smooth graph functions with applications in recommender systems
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
On the Convergence of Single-Timescale Actor-Critic
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Adaptive Few-Shot Learning (AFSL): Tackling Data Scarcity with Stability, Robustness, and Versatility
by: Agrawal, Rishabh
Published: (2025)
by: Agrawal, Rishabh
Published: (2025)
Amortized In-Context Bayesian Posterior Estimation
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications
by: Agrawal, Naman
Published: (2025)
by: Agrawal, Naman
Published: (2025)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
by: Agrawal, Vishakha, et al.
Published: (2025)
by: Agrawal, Vishakha, et al.
Published: (2025)
Illuminate: A novel approach for depression detection with explainable analysis and proactive therapy using prompt engineering
by: Agrawal, Aryan
Published: (2024)
by: Agrawal, Aryan
Published: (2024)
A New Rejection Sampling Approach to $k$-$\mathtt{means}$++ With Improved Trade-Offs
by: Shah, Poojan, et al.
Published: (2025)
by: Shah, Poojan, et al.
Published: (2025)
Reinforcement Learning in MDPs with Information-Ordered Policies
by: Zhang, Zhongjun, et al.
Published: (2025)
by: Zhang, Zhongjun, et al.
Published: (2025)
Disentangling impact of capacity, objective, batchsize, estimators, and step-size on flow VI
by: Agrawal, Abhinav, et al.
Published: (2024)
by: Agrawal, Abhinav, et al.
Published: (2024)
Understanding and mitigating difficulties in posterior predictive evaluation
by: Agrawal, Abhinav, et al.
Published: (2024)
by: Agrawal, Abhinav, et al.
Published: (2024)
FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning
by: Agrawal, Pulkit, et al.
Published: (2025)
by: Agrawal, Pulkit, et al.
Published: (2025)
FGM optimization in complex domains using Gaussian process regression based profile generation algorithm
by: Konda, Chaitanya Kumar, et al.
Published: (2025)
by: Konda, Chaitanya Kumar, et al.
Published: (2025)
Posterior Sampling for Continuing Environments
by: Xu, Wanqiao, et al.
Published: (2022)
by: Xu, Wanqiao, et al.
Published: (2022)
Eventually LIL Regret: Almost Sure $\ln\ln T$ Regret for a sub-Gaussian Mixture on Unbounded Data
by: Agrawal, Shubhada, et al.
Published: (2025)
by: Agrawal, Shubhada, et al.
Published: (2025)
On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds
by: Agrawal, Shubhada, et al.
Published: (2025)
by: Agrawal, Shubhada, et al.
Published: (2025)
Improving Few-Shot Cross-Domain Named Entity Recognition by Instruction Tuning a Word-Embedding based Retrieval Augmented Large Language Model
by: Nandi, Subhadip, et al.
Published: (2024)
by: Nandi, Subhadip, et al.
Published: (2024)
Cover meets Robbins while Betting on Bounded Data: $\ln n$ Regret and Almost Sure $\ln\ln n$ Regret
by: Agrawal, Shubhada, et al.
Published: (2026)
by: Agrawal, Shubhada, et al.
Published: (2026)
Regret Tail Characterization of Optimal Bandit Algorithms with Generic Rewards
by: Panda, Subhodip, et al.
Published: (2026)
by: Panda, Subhodip, et al.
Published: (2026)
Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
Investigating Fouling Efficiency in Football Using Expected Booking (xB) Model
by: Azmat, Adnan, et al.
Published: (2024)
by: Azmat, Adnan, et al.
Published: (2024)
Asymptotic Analysis of Sample-averaged Q-learning
by: Panda, Saunak Kumar, et al.
Published: (2024)
by: Panda, Saunak Kumar, et al.
Published: (2024)
Pareto Set Identification With Posterior Sampling
by: Kone, Cyrille, et al.
Published: (2024)
by: Kone, Cyrille, et al.
Published: (2024)
Gibbs Sampling the Posterior of Neural Networks
by: Piccioli, Giovanni, et al.
Published: (2023)
by: Piccioli, Giovanni, et al.
Published: (2023)
Bayesian Optimization for Dynamic Pricing and Learning
by: Anand, Anush, et al.
Published: (2025)
by: Anand, Anush, et al.
Published: (2025)
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Position: Federated Foundation Language Model Post-Training Should Focus on Open-Source Models
by: Agrawal, Nikita, et al.
Published: (2025)
by: Agrawal, Nikita, et al.
Published: (2025)
Two-level Solar Irradiance Clustering with Season Identification: A Comparative Analysis
by: Agrawal, Roshni, et al.
Published: (2025)
by: Agrawal, Roshni, et al.
Published: (2025)
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024)
by: Pari, Jyothish, et al.
Published: (2024)
VeriGate: Verifier-Gated Step-Level Supervision for GRPO
by: Agrawal, Aakriti, et al.
Published: (2026)
by: Agrawal, Aakriti, et al.
Published: (2026)
AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification
by: Makwe, Ansh, et al.
Published: (2025)
by: Makwe, Ansh, et al.
Published: (2025)
Adaptive Multi-Fidelity Reinforcement Learning for Variance Reduction in Engineering Design Optimization
by: Agrawal, Akash, et al.
Published: (2025)
by: Agrawal, Akash, et al.
Published: (2025)
Adaptive Learning of Design Strategies over Non-Hierarchical Multi-Fidelity Models via Policy Alignment
by: Agrawal, Akash, et al.
Published: (2024)
by: Agrawal, Akash, et al.
Published: (2024)
Addressing LLM Diversity by Infusing Random Concepts
by: Agrawal, Pulin, et al.
Published: (2026)
by: Agrawal, Pulin, et al.
Published: (2026)
Similar Items
-
Optimistic Q-learning for average reward and episodic reinforcement learning
by: Agrawal, Priyank, et al.
Published: (2024) -
Dynamic Pricing and Learning with Long-term Reference Effects
by: Agrawal, Shipra, et al.
Published: (2024) -
A Tractable Online Learning Algorithm for the Multinomial Logit Contextual Bandit
by: Agrawal, Priyank, et al.
Published: (2020) -
Dynamic Pricing and Advertising with Demand Learning
by: Agrawal, Shipra, et al.
Published: (2023) -
Spectral Thompson sampling
by: Kocak, Tomas, et al.
Published: (2026)