Quantizer Design for Finite Model Approximations, Model Learning, and Quantized Q-Learning for MDPs with Unbounded Spaces
Fuente:
arXiv
Guardado en:
| Autores principales: | Bicer, Osman, Kara, Ali D., Yuksel, Serdar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Near Optimal Approximations and Finite Memory Policies for POMPDs with Continuous Spaces
por: Kara, Ali Devran, et al.
Publicado: (2024)
por: Kara, Ali Devran, et al.
Publicado: (2024)
Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning
por: Kara, Ali Devran, et al.
Publicado: (2024)
por: Kara, Ali Devran, et al.
Publicado: (2024)
Robustness to Model Approximation, Model Learning From Data, and Sample Complexity in Wasserstein Regular MDPs
por: Zhou, Yichen, et al.
Publicado: (2024)
por: Zhou, Yichen, et al.
Publicado: (2024)
Q-Learning for Stochastic Control under General Information Structures and Non-Markovian Environments
por: Kara, Ali Devran, et al.
Publicado: (2023)
por: Kara, Ali Devran, et al.
Publicado: (2023)
Kernel Mean Embedding Topology: Weak and Strong Forms for Stochastic Kernels and Implications for Model Learning
por: Saldi, Naci, et al.
Publicado: (2025)
por: Saldi, Naci, et al.
Publicado: (2025)
Reinforcement Learning with Function Approximation for Non-Markov Processes
por: Kara, Ali Devran
Publicado: (2026)
por: Kara, Ali Devran
Publicado: (2026)
Data-Driven Non-Parametric Model Learning and Adaptive Control of MDPs with Borel spaces: Identifiability and Near Optimal Design
por: Mrani-Zentar, Omar, et al.
Publicado: (2025)
por: Mrani-Zentar, Omar, et al.
Publicado: (2025)
Refined Bounds on Near Optimality Finite Window Policies in POMDPs and Their Reinforcement Learning
por: Demirci, Yunus Emre, et al.
Publicado: (2024)
por: Demirci, Yunus Emre, et al.
Publicado: (2024)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
por: Nguyen, Anh Duc, et al.
Publicado: (2025)
por: Nguyen, Anh Duc, et al.
Publicado: (2025)
An Optimal Control Approach To Transformer Training
por: Akman, Kağan, et al.
Publicado: (2026)
por: Akman, Kağan, et al.
Publicado: (2026)
Reinforcement Learning for Discounted and Ergodic Control of Diffusion Processes
por: Bayraktar, Erhan, et al.
Publicado: (2026)
por: Bayraktar, Erhan, et al.
Publicado: (2026)
Online Distributed Learning with Quantized Finite-Time Coordination
por: Bastianello, Nicola, et al.
Publicado: (2023)
por: Bastianello, Nicola, et al.
Publicado: (2023)
Satisficing Paths and Independent Multi-Agent Reinforcement Learning in Stochastic Games
por: Yongacoglu, Bora, et al.
Publicado: (2021)
por: Yongacoglu, Bora, et al.
Publicado: (2021)
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
por: Li, Dongyue, et al.
Publicado: (2026)
por: Li, Dongyue, et al.
Publicado: (2026)
Q-Learning under Finite Model Uncertainty
por: Sester, Julian, et al.
Publicado: (2024)
por: Sester, Julian, et al.
Publicado: (2024)
Learning POMDPs with Linear Function Approximation and Finite Memory
por: Kara, Ali Devran
Publicado: (2025)
por: Kara, Ali Devran
Publicado: (2025)
Planning and Learning in Average Risk-aware MDPs
por: Wang, Weikai, et al.
Publicado: (2025)
por: Wang, Weikai, et al.
Publicado: (2025)
Another Look at Partially Observed Optimal Stochastic Control: Existence, Ergodicity, and Approximations without Belief-Reduction
por: Yüksel, Serdar
Publicado: (2023)
por: Yüksel, Serdar
Publicado: (2023)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
por: Zhang, Zhongjun, et al.
Publicado: (2026)
por: Zhang, Zhongjun, et al.
Publicado: (2026)
Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
por: Kaya, Ege C., et al.
Publicado: (2026)
por: Kaya, Ege C., et al.
Publicado: (2026)
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
por: Chen, Xin, et al.
Publicado: (2024)
por: Chen, Xin, et al.
Publicado: (2024)
Efficient Model-Free Exploration in Low-Rank MDPs
por: Mhammedi, Zakaria, et al.
Publicado: (2023)
por: Mhammedi, Zakaria, et al.
Publicado: (2023)
Stochastic Approximation with Unbounded Markovian Noise: A General-Purpose Theorem
por: Haque, Shaan Ul, et al.
Publicado: (2024)
por: Haque, Shaan Ul, et al.
Publicado: (2024)
Model approximation in MDPs with unbounded per-step cost
por: Bozkurt, Berk, et al.
Publicado: (2024)
por: Bozkurt, Berk, et al.
Publicado: (2024)
PARQ: Piecewise-Affine Regularized Quantization
por: Jin, Lisa, et al.
Publicado: (2025)
por: Jin, Lisa, et al.
Publicado: (2025)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
por: Chae, Woojin, et al.
Publicado: (2024)
por: Chae, Woojin, et al.
Publicado: (2024)
A Simple Finite-Time Analysis of TD Learning with Linear Function Approximation
por: Mitra, Aritra
Publicado: (2024)
por: Mitra, Aritra
Publicado: (2024)
Quantization Avoids Saddle Points in Distributed Optimization
por: Bo, Yanan, et al.
Publicado: (2024)
por: Bo, Yanan, et al.
Publicado: (2024)
Sensitivity of Filter Kernels and Robustness Bounds to Transition and Measurement Kernel Perturbations in Partially Observable Stochastic Control
por: Demirci, Yunus Emre, et al.
Publicado: (2025)
por: Demirci, Yunus Emre, et al.
Publicado: (2025)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
por: Zhang, Runyu, et al.
Publicado: (2023)
por: Zhang, Runyu, et al.
Publicado: (2023)
Incremental Learning of Sparse Attention Patterns in Transformers
por: Yüksel, Oğuz Kaan, et al.
Publicado: (2026)
por: Yüksel, Oğuz Kaan, et al.
Publicado: (2026)
Sliding Window Codes: Near-Optimality and Q-Learning for Zero-Delay Coding
por: Cregg, Liam, et al.
Publicado: (2023)
por: Cregg, Liam, et al.
Publicado: (2023)
Representative Action Selection for Large Action Space: From Bandits to MDPs
por: Zhou, Quan, et al.
Publicado: (2025)
por: Zhou, Quan, et al.
Publicado: (2025)
A Finite-Time Analysis of TD Learning with Linear Function Approximation without Projections or Strong Convexity
por: Lee, Wei-Cheng, et al.
Publicado: (2025)
por: Lee, Wei-Cheng, et al.
Publicado: (2025)
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
por: Zhao, Pengxiang, et al.
Publicado: (2025)
por: Zhao, Pengxiang, et al.
Publicado: (2025)
Decentralized Optimization on Compact Submanifolds by Quantized Riemannian Gradient Tracking
por: Chen, Jun, et al.
Publicado: (2025)
por: Chen, Jun, et al.
Publicado: (2025)
Approximations and Learning for Decentralized Stochastic Control and Near Optimal Finite Window Policies
por: Mrani-Zentar, Omar, et al.
Publicado: (2026)
por: Mrani-Zentar, Omar, et al.
Publicado: (2026)
On Borkar and Young Relaxed Control Topologies and Continuous Dependence of Invariant Measures on Control Policy
por: Yüksel, Serdar
Publicado: (2023)
por: Yüksel, Serdar
Publicado: (2023)
Faster Fixed-Point Methods for Multichain MDPs
por: Zurek, Matthew, et al.
Publicado: (2025)
por: Zurek, Matthew, et al.
Publicado: (2025)
Universal Approximation Theorem for Deep Q-Learning via FBSDE System
por: Qi, Qian
Publicado: (2025)
por: Qi, Qian
Publicado: (2025)
Ejemplares similares
-
Near Optimal Approximations and Finite Memory Policies for POMPDs with Continuous Spaces
por: Kara, Ali Devran, et al.
Publicado: (2024) -
Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning
por: Kara, Ali Devran, et al.
Publicado: (2024) -
Robustness to Model Approximation, Model Learning From Data, and Sample Complexity in Wasserstein Regular MDPs
por: Zhou, Yichen, et al.
Publicado: (2024) -
Q-Learning for Stochastic Control under General Information Structures and Non-Markovian Environments
por: Kara, Ali Devran, et al.
Publicado: (2023) -
Kernel Mean Embedding Topology: Weak and Strong Forms for Stochastic Kernels and Implications for Model Learning
por: Saldi, Naci, et al.
Publicado: (2025)