Solving Richly Constrained Reinforcement Learning through State Augmentation and Reward Penalties
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Hao, Mai, Tien, Varakantham, Pradeep, Hoang, Minh Huy |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement Learning
by: Hoang, Huy, et al.
Published: (2023)
by: Hoang, Huy, et al.
Published: (2023)
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2025)
by: Hoang, Huy, et al.
Published: (2025)
Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
by: Lu, Yuxiao, et al.
Published: (2023)
by: Lu, Yuxiao, et al.
Published: (2023)
Imitating Cost-Constrained Behaviors in Reinforcement Learning
by: Shao, Qian, et al.
Published: (2024)
by: Shao, Qian, et al.
Published: (2024)
On Generalization Across Environments In Multi-Objective Reinforcement Learning
by: Teoh, Jayden, et al.
Published: (2025)
by: Teoh, Jayden, et al.
Published: (2025)
Offline Safe Reinforcement Learning Using Trajectory Classification
by: Gong, Ze, et al.
Published: (2024)
by: Gong, Ze, et al.
Published: (2024)
Regret-Based Defense in Adversarial Reinforcement Learning
by: Belaire, Roman, et al.
Published: (2023)
by: Belaire, Roman, et al.
Published: (2023)
Enhancing the Hierarchical Environment Design via Generative Trajectory Modeling
by: Li, Dexun, et al.
Published: (2023)
by: Li, Dexun, et al.
Published: (2023)
Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning
by: Bui, The Viet, et al.
Published: (2025)
by: Bui, The Viet, et al.
Published: (2025)
Improving Environment Novelty Quantification for Effective Unsupervised Environment Design
by: Teoh, Jayden, et al.
Published: (2025)
by: Teoh, Jayden, et al.
Published: (2025)
Automatic LLM Red Teaming
by: Belaire, Roman, et al.
Published: (2025)
by: Belaire, Roman, et al.
Published: (2025)
On Minimizing Adversarial Counterfactual Error in Adversarial RL
by: Belaire, Roman, et al.
Published: (2024)
by: Belaire, Roman, et al.
Published: (2024)
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression
by: Ge, Zichang, et al.
Published: (2025)
by: Ge, Zichang, et al.
Published: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment
by: Vamplew, Peter, et al.
Published: (2026)
by: Vamplew, Peter, et al.
Published: (2026)
Stochastic Constrained Decentralized Optimization for Machine Learning with Fewer Data Oracles: a Gradient Sliding Approach
by: Nguyen, Hoang Huy, et al.
Published: (2024)
by: Nguyen, Hoang Huy, et al.
Published: (2024)
Misclassification excess risk bounds for 1-bit matrix completion
by: Mai, The Tien
Published: (2023)
by: Mai, The Tien
Published: (2023)
Concentration properties of fractional posterior in 1-bit matrix completion
by: Mai, The Tien
Published: (2024)
by: Mai, The Tien
Published: (2024)
Misclassification excess risk bounds for PAC-Bayesian classification via convexified loss
by: Mai, The Tien
Published: (2024)
by: Mai, The Tien
Published: (2024)
Stochastic Penalty-Barrier Methods for Constrained Machine Learning
by: Bosák, Adam, et al.
Published: (2026)
by: Bosák, Adam, et al.
Published: (2026)
Proactive Constrained Policy Optimization with Preemptive Penalty
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
Enhanced Penalty-based Bidirectional Reinforcement Learning Algorithms
by: Pula, Sai Gana Sandeep, et al.
Published: (2025)
by: Pula, Sai Gana Sandeep, et al.
Published: (2025)
Heavy Lasso: sparse penalized regression under heavy-tailed noise via data-augmented soft-thresholding
by: Mai, The Tien
Published: (2025)
by: Mai, The Tien
Published: (2025)
Robust low-rank estimation with multiple binary responses using pairwise AUC loss
by: Mai, The Tien
Published: (2026)
by: Mai, The Tien
Published: (2026)
Exponential Lasso: robust sparse penalization under heavy-tailed noise and outliers with exponential-type loss
by: Mai, The Tien
Published: (2025)
by: Mai, The Tien
Published: (2025)
Variable-Agnostic Causal Exploration for Reinforcement Learning
by: Nguyen, Minh Hoang, et al.
Published: (2024)
by: Nguyen, Minh Hoang, et al.
Published: (2024)
Reinforcement Learning with Exogenous States and Rewards
by: Trimponias, George, et al.
Published: (2023)
by: Trimponias, George, et al.
Published: (2023)
ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift Regularization
by: Bui, The Viet, et al.
Published: (2024)
by: Bui, The Viet, et al.
Published: (2024)
A Lyapunov Drift-Plus-Penalty Method Tailored for Reinforcement Learning with Queue Stability
by: Xu, Wenhan, et al.
Published: (2025)
by: Xu, Wenhan, et al.
Published: (2025)
DynaSearcher: Dynamic Knowledge Graph Augmented Search Agent via Multi-Reward Reinforcement Learning
by: Hao, Chuzhan, et al.
Published: (2025)
by: Hao, Chuzhan, et al.
Published: (2025)
Improving Pareto Set Learning for Expensive Multi-objective Optimization via Stein Variational Hypernetworks
by: Nguyen, Minh-Duc, et al.
Published: (2024)
by: Nguyen, Minh-Duc, et al.
Published: (2024)
Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning
by: Mai, Tan-Ha, et al.
Published: (2025)
by: Mai, Tan-Ha, et al.
Published: (2025)
Self-Supervised Penalty-Based Learning for Robust Constrained Optimization
by: Benslimane, Wyame, et al.
Published: (2025)
by: Benslimane, Wyame, et al.
Published: (2025)
PAC-Bayesian risk bounds for fully connected deep neural network with Gaussian priors
by: Mai, The Tien
Published: (2025)
by: Mai, The Tien
Published: (2025)
Kullback-Leibler excess risk bounds for exponential weighted aggregation in Generalized linear models
by: Mai, The Tien
Published: (2025)
by: Mai, The Tien
Published: (2025)
Adaptive posterior concentration rates for sparse high-dimensional linear regression with random design and unknown error variance
by: Mai, The Tien
Published: (2024)
by: Mai, The Tien
Published: (2024)
Concentration of a sparse Bayesian model with Horseshoe prior in estimating high-dimensional precision matrix
by: Mai, The Tien
Published: (2024)
by: Mai, The Tien
Published: (2024)
Misclassification bounds for PAC-Bayesian sparse deep learning
by: Mai, The Tien
Published: (2024)
by: Mai, The Tien
Published: (2024)
Similar Items
-
Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement Learning
by: Hoang, Huy, et al.
Published: (2023) -
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
by: Hoang, Huy, et al.
Published: (2024) -
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024) -
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2025) -
Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
by: Lu, Yuxiao, et al.
Published: (2023)