From Relative Entropy to Minimax: A Unified Framework for Coverage in MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Xihe, Mitra, Urbashi, Javidi, Tara |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$κ$-Explorer: A Unified Framework for Active Model Estimation in MDPs
by: Gu, Xihe, et al.
Published: (2026)
by: Gu, Xihe, et al.
Published: (2026)
Coverage Analysis for Digital Cousin Selection -- Improving Multi-Environment Q-Learning
by: Bozkus, Talha, et al.
Published: (2024)
by: Bozkus, Talha, et al.
Published: (2024)
Coverage Analysis of Multi-Environment Q-Learning Algorithms for Wireless Network Optimization
by: Bozkus, Talha, et al.
Published: (2024)
by: Bozkus, Talha, et al.
Published: (2024)
Trojan Cleansing with Neural Collapse
by: Gu, Xihe, et al.
Published: (2024)
by: Gu, Xihe, et al.
Published: (2024)
A Multi-Agent Multi-Environment Mixed Q-Learning for Partially Decentralized Wireless Network Optimization
by: Bozkus, Talha, et al.
Published: (2024)
by: Bozkus, Talha, et al.
Published: (2024)
Multi-Timescale Ensemble Q-learning for Markov Decision Process Policy Optimization
by: Bozkus, Talha, et al.
Published: (2024)
by: Bozkus, Talha, et al.
Published: (2024)
Partially Decentralized Multi-Agent Q-Learning via Digital Cousins for Wireless Networks
by: Bozkus, Talha, et al.
Published: (2025)
by: Bozkus, Talha, et al.
Published: (2025)
Consequences of Kernel Regularity for Bandit Optimization
by: Lee, Madison, et al.
Published: (2025)
by: Lee, Madison, et al.
Published: (2025)
Leveraging Digital Cousins for Ensemble Q-Learning in Large-Scale Wireless Networks
by: Bozkus, Talha, et al.
Published: (2024)
by: Bozkus, Talha, et al.
Published: (2024)
Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing
by: Soleymani, Mahdi, et al.
Published: (2025)
by: Soleymani, Mahdi, et al.
Published: (2025)
ModShift: Model Privacy via Designed Shifts
by: Kherani, Nomaan A., et al.
Published: (2025)
by: Kherani, Nomaan A., et al.
Published: (2025)
Asymmetric Graph Error Control with Low Complexity in Causal Bandits
by: Peng, Chen, et al.
Published: (2024)
by: Peng, Chen, et al.
Published: (2024)
Convex and Non-convex Federated Learning with Stale Stochastic Gradients: Diminishing Step Size is All You Need
by: Zheng, Xinran, et al.
Published: (2026)
by: Zheng, Xinran, et al.
Published: (2026)
Minimax Generalized Cross-Entropy
by: Bondugula, Kartheek, et al.
Published: (2026)
by: Bondugula, Kartheek, et al.
Published: (2026)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
by: Boudart, Pierre, et al.
Published: (2026)
by: Boudart, Pierre, et al.
Published: (2026)
CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning
by: Javadi, Amirhosein, et al.
Published: (2026)
by: Javadi, Amirhosein, et al.
Published: (2026)
Primal Dual Continual Learning: Balancing Stability and Plasticity through Adaptive Memory Allocation
by: Elenter, Juan, et al.
Published: (2023)
by: Elenter, Juan, et al.
Published: (2023)
Semi-supervised Domain Adaptation on Graphs with Contrastive Learning and Minimax Entropy
by: Xiao, Jiaren, et al.
Published: (2023)
by: Xiao, Jiaren, et al.
Published: (2023)
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
by: Lee, Joongkyu, et al.
Published: (2024)
by: Lee, Joongkyu, et al.
Published: (2024)
Evolutionary Multi-Task Optimization for LLM-Guided Program Discovery
by: Gozeten, Halil Alperen, et al.
Published: (2026)
by: Gozeten, Halil Alperen, et al.
Published: (2026)
Learning from Similar Linear Representations: Adaptivity, Minimaxity, and Robustness
by: Tian, Ye, et al.
Published: (2023)
by: Tian, Ye, et al.
Published: (2023)
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Reward Guidance for Reinforcement Learning Tasks Based on Large Language Models: The LMGT Framework
by: Deng, Yongxin, et al.
Published: (2024)
by: Deng, Yongxin, et al.
Published: (2024)
Effective Sparsity: A Unified Framework via Normalized Entropy and the Effective Number of Nonzeros
by: He, Haoyu, et al.
Published: (2026)
by: He, Haoyu, et al.
Published: (2026)
Time-Constrained Robust MDPs
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
A Unified Framework for Entropy Search and Expected Improvement in Bayesian Optimization
by: Cheng, Nuojin, et al.
Published: (2025)
by: Cheng, Nuojin, et al.
Published: (2025)
SureFED: Robust Federated Learning via Uncertainty-Aware Inward and Outward Inspection
by: Heydaribeni, Nasimeh, et al.
Published: (2023)
by: Heydaribeni, Nasimeh, et al.
Published: (2023)
Meta-TTT: A Meta-learning Minimax Framework For Test-Time Training
by: Tao, Chen, et al.
Published: (2024)
by: Tao, Chen, et al.
Published: (2024)
Best-of-Majority: Minimax-Optimal Strategy for Pass@$k$ Inference Scaling
by: Di, Qiwei, et al.
Published: (2025)
by: Di, Qiwei, et al.
Published: (2025)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Zeroth-Order Stochastic Mirror Descent Algorithms for Minimax Excess Risk Optimization
by: Gu, Zhihao, et al.
Published: (2024)
by: Gu, Zhihao, et al.
Published: (2024)
Robust Parameter Learning for Uncertain MDPs
by: Schnitzer, Yannik, et al.
Published: (2026)
by: Schnitzer, Yannik, et al.
Published: (2026)
On Value Iteration Convergence in Connected MDPs
by: Mustafin, Arsenii, et al.
Published: (2024)
by: Mustafin, Arsenii, et al.
Published: (2024)
MDPs with a State Sensing Cost
by: Kapoor, Vansh, et al.
Published: (2025)
by: Kapoor, Vansh, et al.
Published: (2025)
Truly No-Regret Learning in Constrained MDPs
by: Müller, Adrian, et al.
Published: (2024)
by: Müller, Adrian, et al.
Published: (2024)
A Unifying View of Coverage in Linear Off-Policy Evaluation
by: Amortila, Philip, et al.
Published: (2026)
by: Amortila, Philip, et al.
Published: (2026)
Similar Items
-
$κ$-Explorer: A Unified Framework for Active Model Estimation in MDPs
by: Gu, Xihe, et al.
Published: (2026) -
Coverage Analysis for Digital Cousin Selection -- Improving Multi-Environment Q-Learning
by: Bozkus, Talha, et al.
Published: (2024) -
Coverage Analysis of Multi-Environment Q-Learning Algorithms for Wireless Network Optimization
by: Bozkus, Talha, et al.
Published: (2024) -
Trojan Cleansing with Neural Collapse
by: Gu, Xihe, et al.
Published: (2024) -
A Multi-Agent Multi-Environment Mixed Q-Learning for Partially Decentralized Wireless Network Optimization
by: Bozkus, Talha, et al.
Published: (2024)