Towards Optimal Adversarial Robust Q-learning with Bellman Infinity-error
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haoran, Zhang, Zicheng, Luo, Wang, Han, Congying, Hu, Yudong, Guo, Tiande, Liao, Shichen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Optimal Adversarial Robust Reinforcement Learning with Infinity Measurement Error
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
On the Tension Between Optimality and Adversarial Robustness in Policy Optimization
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
Dual Alignment Maximin Optimization for Offline Model-based RL
von: Zhou, Chi, et al.
Veröffentlicht: (2025)
von: Zhou, Chi, et al.
Veröffentlicht: (2025)
Mitigating Distribution Shift in Model-based Offline RL via Shifts-aware Reward Learning
von: Luo, Wang, et al.
Veröffentlicht: (2024)
von: Luo, Wang, et al.
Veröffentlicht: (2024)
Purity Law for Generalizable Neural TSP Solvers
von: Liu, Wenzhao, et al.
Veröffentlicht: (2025)
von: Liu, Wenzhao, et al.
Veröffentlicht: (2025)
Robust Accelerated Adaptive Search: High-Probability Complexity Bounds under Bounded-Moment Stochastic Oracles
von: Zhang, Shunzhi, et al.
Veröffentlicht: (2026)
von: Zhang, Shunzhi, et al.
Veröffentlicht: (2026)
A Fast Anti-Jamming Cognitive Radar Deployment Algorithm Based on Reinforcement Learning
von: Cai, Wencheng, et al.
Veröffentlicht: (2025)
von: Cai, Wencheng, et al.
Veröffentlicht: (2025)
Understanding Oversmoothing in Diffusion-Based GNNs From the Perspective of Operator Semigroup Theory
von: Zhao, Weichen, et al.
Veröffentlicht: (2024)
von: Zhao, Weichen, et al.
Veröffentlicht: (2024)
Applying Opponent Modeling for Automatic Bidding in Online Repeated Auctions
von: Hu, Yudong, et al.
Veröffentlicht: (2022)
von: Hu, Yudong, et al.
Veröffentlicht: (2022)
A-PSRO: A Unified Strategy Learning Method with Advantage Function for Normal-form Games
von: Hu, Yudong, et al.
Veröffentlicht: (2023)
von: Hu, Yudong, et al.
Veröffentlicht: (2023)
A Near-optimal, Scalable and Parallelizable Framework for Stochastic Bandits Robust to Adversarial Corruptions and Beyond
von: Hu, Zicheng, et al.
Veröffentlicht: (2025)
von: Hu, Zicheng, et al.
Veröffentlicht: (2025)
On the optimal pivot path of simplex method for linear programming based on reinforcement learning
von: Li, Anqi, et al.
Veröffentlicht: (2022)
von: Li, Anqi, et al.
Veröffentlicht: (2022)
Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2025)
Towards Blackwell Optimality: Bellman Optimality Is All You Can Get
von: Boone, Victor, et al.
Veröffentlicht: (2025)
von: Boone, Victor, et al.
Veröffentlicht: (2025)
Preference-based opponent shaping in differentiable games
von: Qiao, Xinyu, et al.
Veröffentlicht: (2024)
von: Qiao, Xinyu, et al.
Veröffentlicht: (2024)
DR-BFR: Degradation Representation with Diffusion Models for Blind Face Restoration
von: Qiu, Xinmin, et al.
Veröffentlicht: (2024)
von: Qiu, Xinmin, et al.
Veröffentlicht: (2024)
ShiQ: Bringing back Bellman to LLMs
von: Clavier, Pierre, et al.
Veröffentlicht: (2025)
von: Clavier, Pierre, et al.
Veröffentlicht: (2025)
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
Resolving Endpoint Underfitting in Diffusion Bridges via Noise Alignment
von: Gao, Yurong, et al.
Veröffentlicht: (2026)
von: Gao, Yurong, et al.
Veröffentlicht: (2026)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Role of Bellman Constraints
von: Xu, Tian, et al.
Veröffentlicht: (2026)
von: Xu, Tian, et al.
Veröffentlicht: (2026)
Hierarchical Refinement: Optimal Transport to Infinity and Beyond
von: Halmos, Peter, et al.
Veröffentlicht: (2025)
von: Halmos, Peter, et al.
Veröffentlicht: (2025)
Relating Checkpoint Update Probabilities to Momentum Parameters in Single-Loop Variance Reduction Methods
von: Liu, Hai, et al.
Veröffentlicht: (2026)
von: Liu, Hai, et al.
Veröffentlicht: (2026)
Fitted $Q$ Evaluation Without Bellman Completeness via Stationary Weighting
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
von: van der Laan, Lars, et al.
Veröffentlicht: (2025)
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
von: Omura, Motoki, et al.
Veröffentlicht: (2025)
von: Omura, Motoki, et al.
Veröffentlicht: (2025)
Robust Decentralized Multi-armed Bandits: From Corruption-Resilience to Byzantine-Resilience
von: Hu, Zicheng, et al.
Veröffentlicht: (2025)
von: Hu, Zicheng, et al.
Veröffentlicht: (2025)
Certifiably Safe Manipulation of Deformable Linear Objects via Joint Shape and Tension Prediction
von: Zhang, Yiting, et al.
Veröffentlicht: (2025)
von: Zhang, Yiting, et al.
Veröffentlicht: (2025)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
von: Chen, Lu, et al.
Veröffentlicht: (2025)
von: Chen, Lu, et al.
Veröffentlicht: (2025)
StyO: Stylize Your Face in Only One-shot
von: Li, Bonan, et al.
Veröffentlicht: (2023)
von: Li, Bonan, et al.
Veröffentlicht: (2023)
Regularized Q-learning
von: Lim, Han-Dong, et al.
Veröffentlicht: (2022)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2022)
Toward Evaluating Robustness of Reinforcement Learning with Adversarial Policy
von: Zheng, Xiang, et al.
Veröffentlicht: (2023)
von: Zheng, Xiang, et al.
Veröffentlicht: (2023)
Bellman operator convergence enhancements in reinforcement learning algorithms
von: Kadurha, David Krame, et al.
Veröffentlicht: (2025)
von: Kadurha, David Krame, et al.
Veröffentlicht: (2025)
Towards Fair Class-wise Robustness: Class Optimal Distribution Adversarial Training
von: Zhi, Hongxin, et al.
Veröffentlicht: (2025)
von: Zhi, Hongxin, et al.
Veröffentlicht: (2025)
Multi-Modal Data Fusion for Moisture Content Prediction in Apple Drying
von: Li, Shichen, et al.
Veröffentlicht: (2025)
von: Li, Shichen, et al.
Veröffentlicht: (2025)
BlazeBVD: Make Scale-Time Equalization Great Again for Blind Video Deflickering
von: Qiu, Xinmin, et al.
Veröffentlicht: (2024)
von: Qiu, Xinmin, et al.
Veröffentlicht: (2024)
Bellman Error Centering
von: Chen, Xingguo, et al.
Veröffentlicht: (2025)
von: Chen, Xingguo, et al.
Veröffentlicht: (2025)
Bellman Optimal Stepsize Straightening of Flow-Matching Models
von: Nguyen, Bao, et al.
Veröffentlicht: (2023)
von: Nguyen, Bao, et al.
Veröffentlicht: (2023)
Coupled VAE: Improved Accuracy and Robustness of a Variational Autoencoder
von: Cao, Shichen, et al.
Veröffentlicht: (2019)
von: Cao, Shichen, et al.
Veröffentlicht: (2019)
Optimal Transport Regularized Divergences: Application to Adversarial Robustness
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2023)
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2023)
A Finite Sample Complexity Bound for Distributionally Robust Q-learning
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Towards Optimal Adversarial Robust Reinforcement Learning with Infinity Measurement Error
von: Li, Haoran, et al.
Veröffentlicht: (2025) -
On the Tension Between Optimality and Adversarial Robustness in Policy Optimization
von: Li, Haoran, et al.
Veröffentlicht: (2025) -
Dual Alignment Maximin Optimization for Offline Model-based RL
von: Zhou, Chi, et al.
Veröffentlicht: (2025) -
Mitigating Distribution Shift in Model-based Offline RL via Shifts-aware Reward Learning
von: Luo, Wang, et al.
Veröffentlicht: (2024) -
Purity Law for Generalizable Neural TSP Solvers
von: Liu, Wenzhao, et al.
Veröffentlicht: (2025)