Robust Regularized Policy Iteration under Transition Uncertainty
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Hongqiang, Fu, Zhenghui, Tang, Weihao, Wang, Pengfei, Sun, Yiding, Huang, Qixian, Zhang, Dongxu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning
von: Huang, Qixian, et al.
Veröffentlicht: (2026)
von: Huang, Qixian, et al.
Veröffentlicht: (2026)
Offline Policy Optimization with Posterior Sampling
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
von: Eshwar, S. R., et al.
Veröffentlicht: (2025)
von: Eshwar, S. R., et al.
Veröffentlicht: (2025)
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
von: Sun, Cheng, et al.
Veröffentlicht: (2025)
von: Sun, Cheng, et al.
Veröffentlicht: (2025)
Robust Uncertainty Estimation under Distribution Shift via Difference Reconstruction
von: Xu, Xinran, et al.
Veröffentlicht: (2026)
von: Xu, Xinran, et al.
Veröffentlicht: (2026)
Relative Policy-Transition Optimization for Fast Policy Transfer
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
Are LLMs Better GNN Helpers? Rethinking Robust Graph Learning under Deficiencies with Iterative Refinement
von: Wang, Zhaoyan, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyan, et al.
Veröffentlicht: (2025)
Uncertainty Regularized Evidential Regression
von: Ye, Kai, et al.
Veröffentlicht: (2024)
von: Ye, Kai, et al.
Veröffentlicht: (2024)
Policy Regularized Distributionally Robust Markov Decision Processes with Linear Function Approximation
von: Gu, Jingwen, et al.
Veröffentlicht: (2025)
von: Gu, Jingwen, et al.
Veröffentlicht: (2025)
Transformers Learn Robust In-Context Regression under Distributional Uncertainty
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
UCPO: Uncertainty-Aware Policy Optimization
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
Avoiding $\mathbf{exp(R_{max})}$ scaling in RLHF through Preference-based Exploration
von: Chen, Mingyu, et al.
Veröffentlicht: (2025)
von: Chen, Mingyu, et al.
Veröffentlicht: (2025)
COPR: Continual Human Preference Learning via Optimal Policy Regularization
von: Zhang, Han, et al.
Veröffentlicht: (2024)
von: Zhang, Han, et al.
Veröffentlicht: (2024)
Optimistic Policy Regularization
von: Pham, Mai, et al.
Veröffentlicht: (2026)
von: Pham, Mai, et al.
Veröffentlicht: (2026)
IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning
von: Qin, Yihao, et al.
Veröffentlicht: (2026)
von: Qin, Yihao, et al.
Veröffentlicht: (2026)
REG: A Regularization Optimizer for Robust Training Dynamics
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
Robust Bidirectional Associative Memory via Regularization Inspired by the Subspace Rotation Algorithm
von: Lin, Ci, et al.
Veröffentlicht: (2025)
von: Lin, Ci, et al.
Veröffentlicht: (2025)
The Impact of Machine Learning Uncertainty on the Robustness of Counterfactual Explanations
von: Christodoulou, Leonidas, et al.
Veröffentlicht: (2026)
von: Christodoulou, Leonidas, et al.
Veröffentlicht: (2026)
Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data Corruptions
von: Yang, Rui, et al.
Veröffentlicht: (2024)
von: Yang, Rui, et al.
Veröffentlicht: (2024)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
von: He, Longxiang, et al.
Veröffentlicht: (2025)
von: He, Longxiang, et al.
Veröffentlicht: (2025)
IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck
von: Deng, Huilin, et al.
Veröffentlicht: (2026)
von: Deng, Huilin, et al.
Veröffentlicht: (2026)
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
von: Wen, Muning, et al.
Veröffentlicht: (2024)
von: Wen, Muning, et al.
Veröffentlicht: (2024)
Adaptive Conformal Guidance for Learning under Uncertainty
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Solution-oriented Agent-based Models Generation with Verifier-assisted Iterative In-context Learning
von: Niu, Tong, et al.
Veröffentlicht: (2024)
von: Niu, Tong, et al.
Veröffentlicht: (2024)
Complexity-Regularized Proximal Policy Optimization
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
Symmetric Behavior Regularized Policy Optimization
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
von: Wang, Lu, et al.
Veröffentlicht: (2024)
von: Wang, Lu, et al.
Veröffentlicht: (2024)
Monte Carlo Tree Search in the Presence of Transition Uncertainty
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2023)
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2023)
Fairness-Regularized Online Optimization with Switching Costs
von: Li, Pengfei, et al.
Veröffentlicht: (2025)
von: Li, Pengfei, et al.
Veröffentlicht: (2025)
Selective Learning: Towards Robust Calibration with Dynamic Regularization
von: Han, Zongbo, et al.
Veröffentlicht: (2024)
von: Han, Zongbo, et al.
Veröffentlicht: (2024)
Active Asymmetric Multi-Agent Multimodal Learning under Uncertainty
von: Liu, Rui, et al.
Veröffentlicht: (2026)
von: Liu, Rui, et al.
Veröffentlicht: (2026)
Learning Scenario Reduction for Two-Stage Robust Optimization with Discrete Uncertainty
von: Lin, Tianjue, et al.
Veröffentlicht: (2026)
von: Lin, Tianjue, et al.
Veröffentlicht: (2026)
Learning Displacement-Robust Representations for Landslide Early Warning under Rainfall Forecast Uncertainty
von: Ozeki, Ren, et al.
Veröffentlicht: (2026)
von: Ozeki, Ren, et al.
Veröffentlicht: (2026)
Advancing Few-Shot Pediatric Arrhythmia Classification with a Novel Contrastive Loss and Multimodal Learning
von: Chen, Yiqiao, et al.
Veröffentlicht: (2025)
von: Chen, Yiqiao, et al.
Veröffentlicht: (2025)
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
von: Na, Byeonghu, et al.
Veröffentlicht: (2026)
von: Na, Byeonghu, et al.
Veröffentlicht: (2026)
Robust Offline Reinforcement Learning with Linearly Structured f-Divergence Regularization
von: Tang, Cheng, et al.
Veröffentlicht: (2024)
von: Tang, Cheng, et al.
Veröffentlicht: (2024)
Robust Causal Discovery under Imperfect Structural Constraints
von: Wang, Zidong, et al.
Veröffentlicht: (2025)
von: Wang, Zidong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026) -
CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning
von: Huang, Qixian, et al.
Veröffentlicht: (2026) -
Offline Policy Optimization with Posterior Sampling
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026) -
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
von: Eshwar, S. R., et al.
Veröffentlicht: (2025) -
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)