Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhattacharyya, Riddhiman, Chakrabarty, Sayak, Banerjee, Imon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
von: Banerjee, Imon, et al.
Veröffentlicht: (2025)
von: Banerjee, Imon, et al.
Veröffentlicht: (2025)
Offline Estimation of Controlled Markov Chains: Minimaxity and Sample Complexity
von: Banerjee, Imon, et al.
Veröffentlicht: (2022)
von: Banerjee, Imon, et al.
Veröffentlicht: (2022)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
von: Qian, Jian, et al.
Veröffentlicht: (2024)
von: Qian, Jian, et al.
Veröffentlicht: (2024)
Time-Constrained Recommendations: Reinforcement Learning Strategies for E-Commerce
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2025)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
Is Sliding Window All You Need? An Open Framework for Long-Sequence Recommendation
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2026)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2026)
PixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2026)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2026)
Why DPO is a Misspecified Estimator and How to Fix It
von: Gopalan, Aditya, et al.
Veröffentlicht: (2025)
von: Gopalan, Aditya, et al.
Veröffentlicht: (2025)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
Eluder-based Regret for Stochastic Contextual MDPs
von: Levy, Orin, et al.
Veröffentlicht: (2022)
von: Levy, Orin, et al.
Veröffentlicht: (2022)
Sample Complexity Characterization for Linear Contextual MDPs
von: Deng, Junze, et al.
Veröffentlicht: (2024)
von: Deng, Junze, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
Stationarity without mean reversion in improper Gaussian processes
von: Ambrogioni, Luca
Veröffentlicht: (2023)
von: Ambrogioni, Luca
Veröffentlicht: (2023)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
von: Yun, Zeyu, et al.
Veröffentlicht: (2023)
von: Yun, Zeyu, et al.
Veröffentlicht: (2023)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
von: Lee, Kyungbok, et al.
Veröffentlicht: (2026)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2026)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
von: Weltevrede, Max, et al.
Veröffentlicht: (2024)
von: Weltevrede, Max, et al.
Veröffentlicht: (2024)
Group-Sensitive Offline Contextual Bandits
von: Guo, Yihong, et al.
Veröffentlicht: (2025)
von: Guo, Yihong, et al.
Veröffentlicht: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
von: Viano, Luca, et al.
Veröffentlicht: (2024)
von: Viano, Luca, et al.
Veröffentlicht: (2024)
On the System Theoretic Offline Learning of Continuous-Time LQR with Exogenous Disturbances
von: Mukherjee, Sayak, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sayak, et al.
Veröffentlicht: (2025)
Near-Optimal Sample Complexity for Online Constrained MDPs
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
von: Shrestha, Aayam, et al.
Veröffentlicht: (2020)
von: Shrestha, Aayam, et al.
Veröffentlicht: (2020)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2023)
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2023)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
Offline Bayesian Aleatoric and Epistemic Uncertainty Quantification and Posterior Value Optimisation in Finite-State MDPs
von: Valdettaro, Filippo, et al.
Veröffentlicht: (2024)
von: Valdettaro, Filippo, et al.
Veröffentlicht: (2024)
Offline Contextual Bandits in the Presence of New Actions
von: Kishimoto, Ren, et al.
Veröffentlicht: (2026)
von: Kishimoto, Ren, et al.
Veröffentlicht: (2026)
Contextual Online Pricing with (Biased) Offline Data
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
Offline Contextual Bandit with Counterfactual Sample Identification
von: Gilotte, Alexandre, et al.
Veröffentlicht: (2025)
von: Gilotte, Alexandre, et al.
Veröffentlicht: (2025)
CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings
von: Bhatia, Randeep, et al.
Veröffentlicht: (2025)
von: Bhatia, Randeep, et al.
Veröffentlicht: (2025)
Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence
von: Lee, Donghwan, et al.
Veröffentlicht: (2026)
von: Lee, Donghwan, et al.
Veröffentlicht: (2026)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs
von: John, Philips George, et al.
Veröffentlicht: (2024)
von: John, Philips George, et al.
Veröffentlicht: (2024)
Entropy-Guided Sampling of Flat Modes in Discrete Spaces
von: Mohanty, Pinaki, et al.
Veröffentlicht: (2025)
von: Mohanty, Pinaki, et al.
Veröffentlicht: (2025)
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
von: Tariq, Amara, et al.
Veröffentlicht: (2025)
von: Tariq, Amara, et al.
Veröffentlicht: (2025)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
von: Wei, Yukuan, et al.
Veröffentlicht: (2025)
von: Wei, Yukuan, et al.
Veröffentlicht: (2025)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
von: Han, Zean, et al.
Veröffentlicht: (2026)
von: Han, Zean, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
von: Banerjee, Imon, et al.
Veröffentlicht: (2025) -
Offline Estimation of Controlled Markov Chains: Minimaxity and Sample Complexity
von: Banerjee, Imon, et al.
Veröffentlicht: (2022) -
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026) -
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
von: Qian, Jian, et al.
Veröffentlicht: (2024) -
Time-Constrained Recommendations: Reinforcement Learning Strategies for E-Commerce
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2025)