Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
Fuente:
arXiv
Guardado en:
| Autores principales: | Bhattacharyya, Riddhiman, Chakrabarty, Sayak, Banerjee, Imon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
por: Banerjee, Imon, et al.
Publicado: (2025)
por: Banerjee, Imon, et al.
Publicado: (2025)
Offline Estimation of Controlled Markov Chains: Minimaxity and Sample Complexity
por: Banerjee, Imon, et al.
Publicado: (2022)
por: Banerjee, Imon, et al.
Publicado: (2022)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
por: Levy, Orin, et al.
Publicado: (2026)
por: Levy, Orin, et al.
Publicado: (2026)
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
por: Qian, Jian, et al.
Publicado: (2024)
por: Qian, Jian, et al.
Publicado: (2024)
Time-Constrained Recommendations: Reinforcement Learning Strategies for E-Commerce
por: Chakrabarty, Sayak, et al.
Publicado: (2025)
por: Chakrabarty, Sayak, et al.
Publicado: (2025)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
por: Chakrabarty, Sayak, et al.
Publicado: (2024)
por: Chakrabarty, Sayak, et al.
Publicado: (2024)
Is Sliding Window All You Need? An Open Framework for Long-Sequence Recommendation
por: Chakrabarty, Sayak, et al.
Publicado: (2026)
por: Chakrabarty, Sayak, et al.
Publicado: (2026)
PixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation
por: Chakrabarty, Sayak, et al.
Publicado: (2026)
por: Chakrabarty, Sayak, et al.
Publicado: (2026)
Why DPO is a Misspecified Estimator and How to Fix It
por: Gopalan, Aditya, et al.
Publicado: (2025)
por: Gopalan, Aditya, et al.
Publicado: (2025)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
por: Chakrabarty, Sayak, et al.
Publicado: (2024)
por: Chakrabarty, Sayak, et al.
Publicado: (2024)
Eluder-based Regret for Stochastic Contextual MDPs
por: Levy, Orin, et al.
Publicado: (2022)
por: Levy, Orin, et al.
Publicado: (2022)
Sample Complexity Characterization for Linear Contextual MDPs
por: Deng, Junze, et al.
Publicado: (2024)
por: Deng, Junze, et al.
Publicado: (2024)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
por: Ackermann, Johannes, et al.
Publicado: (2024)
por: Ackermann, Johannes, et al.
Publicado: (2024)
Stationarity without mean reversion in improper Gaussian processes
por: Ambrogioni, Luca
Publicado: (2023)
por: Ambrogioni, Luca
Publicado: (2023)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
por: Zhang, Zhongjun, et al.
Publicado: (2026)
por: Zhang, Zhongjun, et al.
Publicado: (2026)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
por: Hong, Kihyuk, et al.
Publicado: (2024)
por: Hong, Kihyuk, et al.
Publicado: (2024)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
por: Yun, Zeyu, et al.
Publicado: (2023)
por: Yun, Zeyu, et al.
Publicado: (2023)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
por: Lee, Kyungbok, et al.
Publicado: (2026)
por: Lee, Kyungbok, et al.
Publicado: (2026)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
por: Weltevrede, Max, et al.
Publicado: (2024)
por: Weltevrede, Max, et al.
Publicado: (2024)
Group-Sensitive Offline Contextual Bandits
por: Guo, Yihong, et al.
Publicado: (2025)
por: Guo, Yihong, et al.
Publicado: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024)
por: Viano, Luca, et al.
Publicado: (2024)
On the System Theoretic Offline Learning of Continuous-Time LQR with Exogenous Disturbances
por: Mukherjee, Sayak, et al.
Publicado: (2025)
por: Mukherjee, Sayak, et al.
Publicado: (2025)
Near-Optimal Sample Complexity for Online Constrained MDPs
por: Liu, Chang, et al.
Publicado: (2026)
por: Liu, Chang, et al.
Publicado: (2026)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
por: Shrestha, Aayam, et al.
Publicado: (2020)
por: Shrestha, Aayam, et al.
Publicado: (2020)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
por: Banerjee, Debangshu, et al.
Publicado: (2023)
por: Banerjee, Debangshu, et al.
Publicado: (2023)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
por: Moulin, Antoine, et al.
Publicado: (2025)
por: Moulin, Antoine, et al.
Publicado: (2025)
Offline Bayesian Aleatoric and Epistemic Uncertainty Quantification and Posterior Value Optimisation in Finite-State MDPs
por: Valdettaro, Filippo, et al.
Publicado: (2024)
por: Valdettaro, Filippo, et al.
Publicado: (2024)
Offline Contextual Bandits in the Presence of New Actions
por: Kishimoto, Ren, et al.
Publicado: (2026)
por: Kishimoto, Ren, et al.
Publicado: (2026)
Contextual Online Pricing with (Biased) Offline Data
por: Zhang, Yixuan, et al.
Publicado: (2025)
por: Zhang, Yixuan, et al.
Publicado: (2025)
Offline Contextual Bandit with Counterfactual Sample Identification
por: Gilotte, Alexandre, et al.
Publicado: (2025)
por: Gilotte, Alexandre, et al.
Publicado: (2025)
CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings
por: Bhatia, Randeep, et al.
Publicado: (2025)
por: Bhatia, Randeep, et al.
Publicado: (2025)
Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence
por: Lee, Donghwan, et al.
Publicado: (2026)
por: Lee, Donghwan, et al.
Publicado: (2026)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
por: Li, Long-Fei, et al.
Publicado: (2024)
por: Li, Long-Fei, et al.
Publicado: (2024)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
por: Cassel, Asaf, et al.
Publicado: (2024)
por: Cassel, Asaf, et al.
Publicado: (2024)
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs
por: John, Philips George, et al.
Publicado: (2024)
por: John, Philips George, et al.
Publicado: (2024)
Entropy-Guided Sampling of Flat Modes in Discrete Spaces
por: Mohanty, Pinaki, et al.
Publicado: (2025)
por: Mohanty, Pinaki, et al.
Publicado: (2025)
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
por: Tariq, Amara, et al.
Publicado: (2025)
por: Tariq, Amara, et al.
Publicado: (2025)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
por: Wei, Yukuan, et al.
Publicado: (2025)
por: Wei, Yukuan, et al.
Publicado: (2025)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
por: Han, Zean, et al.
Publicado: (2026)
por: Han, Zean, et al.
Publicado: (2026)
Ejemplares similares
-
CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
por: Banerjee, Imon, et al.
Publicado: (2025) -
Offline Estimation of Controlled Markov Chains: Minimaxity and Sample Complexity
por: Banerjee, Imon, et al.
Publicado: (2022) -
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
por: Levy, Orin, et al.
Publicado: (2026) -
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
por: Qian, Jian, et al.
Publicado: (2024) -
Time-Constrained Recommendations: Reinforcement Learning Strategies for E-Commerce
por: Chakrabarty, Sayak, et al.
Publicado: (2025)