A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration
Fuente:
arXiv
Saved in:
| Main Authors: | Haussmann, Manuel, Çelikok, Mustafa Mert, Kandemir, Melih |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributional Active Inference
by: Akgül, Abdullah, et al.
Published: (2026)
by: Akgül, Abdullah, et al.
Published: (2026)
Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies
by: Boock, Magnus Victor, et al.
Published: (2026)
by: Boock, Magnus Victor, et al.
Published: (2026)
Deterministic Uncertainty Propagation for Improved Model-Based Offline Reinforcement Learning
by: Akgül, Abdullah, et al.
Published: (2024)
by: Akgül, Abdullah, et al.
Published: (2024)
Overcoming Non-stationary Dynamics with Evidential Proximal Policy Optimization
by: Akgül, Abdullah, et al.
Published: (2025)
by: Akgül, Abdullah, et al.
Published: (2025)
Adaptive Ensemble Aggregation for Actor-Critics
by: Werge, Nicklas, et al.
Published: (2025)
by: Werge, Nicklas, et al.
Published: (2025)
Deep Exploration with PAC-Bayes
by: Tasdighi, Bahareh, et al.
Published: (2024)
by: Tasdighi, Bahareh, et al.
Published: (2024)
PAC-Bayesian Soft Actor-Critic Learning
by: Tasdighi, Bahareh, et al.
Published: (2023)
by: Tasdighi, Bahareh, et al.
Published: (2023)
Deep Actor-Critics with Tight Risk Certificates
by: Tasdighi, Bahareh, et al.
Published: (2025)
by: Tasdighi, Bahareh, et al.
Published: (2025)
ObjectRL: An Object-Oriented Reinforcement Learning Codebase
by: Baykal, Gulcin, et al.
Published: (2025)
by: Baykal, Gulcin, et al.
Published: (2025)
Social Cooperation in Conversational AI Agents
by: Çelikok, Mustafa Mert, et al.
Published: (2025)
by: Çelikok, Mustafa Mert, et al.
Published: (2025)
On the Complexity of Learning to Cooperate with Populations of Socially Rational Agents
by: Loftin, Robert, et al.
Published: (2024)
by: Loftin, Robert, et al.
Published: (2024)
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Generalized Fitted Q-Iteration with Clustered Data
by: Hu, Liyuan, et al.
Published: (2025)
by: Hu, Liyuan, et al.
Published: (2025)
Inverse Concave-Utility Reinforcement Learning is Inverse Game Theory
by: Çelikok, Mustafa Mert, et al.
Published: (2024)
by: Çelikok, Mustafa Mert, et al.
Published: (2024)
Continual Learning of Multi-modal Dynamics with External Memory
by: Akgül, Abdullah, et al.
Published: (2022)
by: Akgül, Abdullah, et al.
Published: (2022)
Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning
by: Manda, Kausthubh, et al.
Published: (2025)
by: Manda, Kausthubh, et al.
Published: (2025)
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
by: Suau, Miguel, et al.
Published: (2022)
by: Suau, Miguel, et al.
Published: (2022)
Disentanglement with Factor Quantized Variational Autoencoders
by: Baykal, Gulcin, et al.
Published: (2024)
by: Baykal, Gulcin, et al.
Published: (2024)
EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders
by: Baykal, Gulcin, et al.
Published: (2023)
by: Baykal, Gulcin, et al.
Published: (2023)
Improved Algorithms for Stochastic Linear Bandits Using Tail Bounds for Martingale Mixtures
by: Flynn, Hamish, et al.
Published: (2023)
by: Flynn, Hamish, et al.
Published: (2023)
Uncoupled Learning of Differential Stackelberg Equilibria with Commitments
by: Loftin, Robert, et al.
Published: (2023)
by: Loftin, Robert, et al.
Published: (2023)
Improving Actor-Critic Training with Steerable Action-Value Approximation Errors
by: Tasdighi, Bahareh, et al.
Published: (2024)
by: Tasdighi, Bahareh, et al.
Published: (2024)
Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits
by: Werge, Nicklas, et al.
Published: (2023)
by: Werge, Nicklas, et al.
Published: (2023)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
by: Baran, Orhun Buğra, et al.
Published: (2026)
by: Baran, Orhun Buğra, et al.
Published: (2026)
Adaptive Inference: Theoretical Limits and Unexplored Opportunities
by: Hor, Soheil, et al.
Published: (2024)
by: Hor, Soheil, et al.
Published: (2024)
Robust Fitted-Q-Evaluation and Iteration under Sequentially Exogenous Unobserved Confounders
by: Bruns-Smith, David, et al.
Published: (2023)
by: Bruns-Smith, David, et al.
Published: (2023)
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024)
by: Oren, Yaniv, et al.
Published: (2024)
Fitted Q-Iteration via Max-Plus-Linear Approximation
by: Liu, Y., et al.
Published: (2024)
by: Liu, Y., et al.
Published: (2024)
Latent variable model for high-dimensional point process with structured missingness
by: Sinelnikov, Maksim, et al.
Published: (2024)
by: Sinelnikov, Maksim, et al.
Published: (2024)
Calibrating Bayesian UNet++ for Sub-Seasonal Forecasting
by: Asan, Busra, et al.
Published: (2024)
by: Asan, Busra, et al.
Published: (2024)
A Finite Sample Complexity Bound for Distributionally Robust Q-learning
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
by: Haussmann, Aden
Published: (2025)
by: Haussmann, Aden
Published: (2025)
Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games
by: Jeong, Narim, et al.
Published: (2026)
by: Jeong, Narim, et al.
Published: (2026)
LogRouter: Adaptive Two-Level LLM Routing for Log Question Answering in Big Data Systems
by: Coskuner, Mert, et al.
Published: (2026)
by: Coskuner, Mert, et al.
Published: (2026)
Exponential Convergence Guarantees for Iterative Markovian Fitting
by: Silveri, Marta Gentiloni, et al.
Published: (2025)
by: Silveri, Marta Gentiloni, et al.
Published: (2025)
A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Latent mixed-effect models for high-dimensional longitudinal data
by: Ong, Priscilla, et al.
Published: (2024)
by: Ong, Priscilla, et al.
Published: (2024)
Exponential convergence rate for Iterative Markovian Fitting
by: Sokolov, Kirill, et al.
Published: (2025)
by: Sokolov, Kirill, et al.
Published: (2025)
Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under $τ$-Mixing
by: Halgryn, Leon, et al.
Published: (2026)
by: Halgryn, Leon, et al.
Published: (2026)
Similar Items
-
Distributional Active Inference
by: Akgül, Abdullah, et al.
Published: (2026) -
Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies
by: Boock, Magnus Victor, et al.
Published: (2026) -
Deterministic Uncertainty Propagation for Improved Model-Based Offline Reinforcement Learning
by: Akgül, Abdullah, et al.
Published: (2024) -
Overcoming Non-stationary Dynamics with Evidential Proximal Policy Optimization
by: Akgül, Abdullah, et al.
Published: (2025) -
Adaptive Ensemble Aggregation for Actor-Critics
by: Werge, Nicklas, et al.
Published: (2025)