A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Haussmann, Manuel, Çelikok, Mustafa Mert, Kandemir, Melih |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distributional Active Inference
von: Akgül, Abdullah, et al.
Veröffentlicht: (2026)
von: Akgül, Abdullah, et al.
Veröffentlicht: (2026)
Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies
von: Boock, Magnus Victor, et al.
Veröffentlicht: (2026)
von: Boock, Magnus Victor, et al.
Veröffentlicht: (2026)
Deterministic Uncertainty Propagation for Improved Model-Based Offline Reinforcement Learning
von: Akgül, Abdullah, et al.
Veröffentlicht: (2024)
von: Akgül, Abdullah, et al.
Veröffentlicht: (2024)
Overcoming Non-stationary Dynamics with Evidential Proximal Policy Optimization
von: Akgül, Abdullah, et al.
Veröffentlicht: (2025)
von: Akgül, Abdullah, et al.
Veröffentlicht: (2025)
Adaptive Ensemble Aggregation for Actor-Critics
von: Werge, Nicklas, et al.
Veröffentlicht: (2025)
von: Werge, Nicklas, et al.
Veröffentlicht: (2025)
Deep Exploration with PAC-Bayes
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2024)
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2024)
PAC-Bayesian Soft Actor-Critic Learning
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2023)
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2023)
Deep Actor-Critics with Tight Risk Certificates
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2025)
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2025)
ObjectRL: An Object-Oriented Reinforcement Learning Codebase
von: Baykal, Gulcin, et al.
Veröffentlicht: (2025)
von: Baykal, Gulcin, et al.
Veröffentlicht: (2025)
Social Cooperation in Conversational AI Agents
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
On the Complexity of Learning to Cooperate with Populations of Socially Rational Agents
von: Loftin, Robert, et al.
Veröffentlicht: (2024)
von: Loftin, Robert, et al.
Veröffentlicht: (2024)
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
Generalized Fitted Q-Iteration with Clustered Data
von: Hu, Liyuan, et al.
Veröffentlicht: (2025)
von: Hu, Liyuan, et al.
Veröffentlicht: (2025)
Inverse Concave-Utility Reinforcement Learning is Inverse Game Theory
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2024)
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2024)
Continual Learning of Multi-modal Dynamics with External Memory
von: Akgül, Abdullah, et al.
Veröffentlicht: (2022)
von: Akgül, Abdullah, et al.
Veröffentlicht: (2022)
Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning
von: Manda, Kausthubh, et al.
Veröffentlicht: (2025)
von: Manda, Kausthubh, et al.
Veröffentlicht: (2025)
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
von: Suau, Miguel, et al.
Veröffentlicht: (2022)
von: Suau, Miguel, et al.
Veröffentlicht: (2022)
Disentanglement with Factor Quantized Variational Autoencoders
von: Baykal, Gulcin, et al.
Veröffentlicht: (2024)
von: Baykal, Gulcin, et al.
Veröffentlicht: (2024)
EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders
von: Baykal, Gulcin, et al.
Veröffentlicht: (2023)
von: Baykal, Gulcin, et al.
Veröffentlicht: (2023)
Improved Algorithms for Stochastic Linear Bandits Using Tail Bounds for Martingale Mixtures
von: Flynn, Hamish, et al.
Veröffentlicht: (2023)
von: Flynn, Hamish, et al.
Veröffentlicht: (2023)
Uncoupled Learning of Differential Stackelberg Equilibria with Commitments
von: Loftin, Robert, et al.
Veröffentlicht: (2023)
von: Loftin, Robert, et al.
Veröffentlicht: (2023)
Improving Actor-Critic Training with Steerable Action-Value Approximation Errors
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2024)
von: Tasdighi, Bahareh, et al.
Veröffentlicht: (2024)
Weighted Sequential Bayesian Inference for Non-Stationary Linear Contextual Bandits
von: Werge, Nicklas, et al.
Veröffentlicht: (2023)
von: Werge, Nicklas, et al.
Veröffentlicht: (2023)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
von: Baran, Orhun Buğra, et al.
Veröffentlicht: (2026)
von: Baran, Orhun Buğra, et al.
Veröffentlicht: (2026)
Adaptive Inference: Theoretical Limits and Unexplored Opportunities
von: Hor, Soheil, et al.
Veröffentlicht: (2024)
von: Hor, Soheil, et al.
Veröffentlicht: (2024)
Robust Fitted-Q-Evaluation and Iteration under Sequentially Exogenous Unobserved Confounders
von: Bruns-Smith, David, et al.
Veröffentlicht: (2023)
von: Bruns-Smith, David, et al.
Veröffentlicht: (2023)
Value Improved Actor Critic Algorithms
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
Fitted Q-Iteration via Max-Plus-Linear Approximation
von: Liu, Y., et al.
Veröffentlicht: (2024)
von: Liu, Y., et al.
Veröffentlicht: (2024)
Latent variable model for high-dimensional point process with structured missingness
von: Sinelnikov, Maksim, et al.
Veröffentlicht: (2024)
von: Sinelnikov, Maksim, et al.
Veröffentlicht: (2024)
Calibrating Bayesian UNet++ for Sub-Seasonal Forecasting
von: Asan, Busra, et al.
Veröffentlicht: (2024)
von: Asan, Busra, et al.
Veröffentlicht: (2024)
A Finite Sample Complexity Bound for Distributionally Robust Q-learning
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
von: Haussmann, Aden
Veröffentlicht: (2025)
von: Haussmann, Aden
Veröffentlicht: (2025)
Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games
von: Jeong, Narim, et al.
Veröffentlicht: (2026)
von: Jeong, Narim, et al.
Veröffentlicht: (2026)
LogRouter: Adaptive Two-Level LLM Routing for Log Question Answering in Big Data Systems
von: Coskuner, Mert, et al.
Veröffentlicht: (2026)
von: Coskuner, Mert, et al.
Veröffentlicht: (2026)
Exponential Convergence Guarantees for Iterative Markovian Fitting
von: Silveri, Marta Gentiloni, et al.
Veröffentlicht: (2025)
von: Silveri, Marta Gentiloni, et al.
Veröffentlicht: (2025)
A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
Latent mixed-effect models for high-dimensional longitudinal data
von: Ong, Priscilla, et al.
Veröffentlicht: (2024)
von: Ong, Priscilla, et al.
Veröffentlicht: (2024)
Exponential convergence rate for Iterative Markovian Fitting
von: Sokolov, Kirill, et al.
Veröffentlicht: (2025)
von: Sokolov, Kirill, et al.
Veröffentlicht: (2025)
Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under $τ$-Mixing
von: Halgryn, Leon, et al.
Veröffentlicht: (2026)
von: Halgryn, Leon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Distributional Active Inference
von: Akgül, Abdullah, et al.
Veröffentlicht: (2026) -
Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies
von: Boock, Magnus Victor, et al.
Veröffentlicht: (2026) -
Deterministic Uncertainty Propagation for Improved Model-Based Offline Reinforcement Learning
von: Akgül, Abdullah, et al.
Veröffentlicht: (2024) -
Overcoming Non-stationary Dynamics with Evidential Proximal Policy Optimization
von: Akgül, Abdullah, et al.
Veröffentlicht: (2025) -
Adaptive Ensemble Aggregation for Actor-Critics
von: Werge, Nicklas, et al.
Veröffentlicht: (2025)