Preserving the Privacy of Reward Functions in MDPs through Deception
Fuente:
arXiv
Saved in:
| Main Authors: | Chirra, Shashank Reddy, Varakantham, Pradeep, Paruchuri, Praveen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safety through feedback in Constrained RL
by: Chirra, Shashank Reddy, et al.
Published: (2024)
by: Chirra, Shashank Reddy, et al.
Published: (2024)
On Discovering Algorithms for Adversarial Imitation Learning
by: Chirra, Shashank Reddy, et al.
Published: (2025)
by: Chirra, Shashank Reddy, et al.
Published: (2025)
Efficient Unsupervised Environment Design through Hierarchical Policy Representation Learning
by: Li, Dexun, et al.
Published: (2026)
by: Li, Dexun, et al.
Published: (2026)
PyVRP$^+$: LLM-Driven Metacognitive Heuristic Evolution for Hybrid Genetic Search in Vehicle Routing Problems
by: Malik, Manuj, et al.
Published: (2026)
by: Malik, Manuj, et al.
Published: (2026)
Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
by: Lu, Yuxiao, et al.
Published: (2023)
by: Lu, Yuxiao, et al.
Published: (2023)
Enhancing the Hierarchical Environment Design via Generative Trajectory Modeling
by: Li, Dexun, et al.
Published: (2023)
by: Li, Dexun, et al.
Published: (2023)
Unlocking Large Language Model's Planning Capabilities with Maximum Diversity Fine-tuning
by: Li, Wenjun, et al.
Published: (2024)
by: Li, Wenjun, et al.
Published: (2024)
Offline Safe Policy Optimization From Heterogeneous Feedback
by: Gong, Ze, et al.
Published: (2025)
by: Gong, Ze, et al.
Published: (2025)
EduQate: Generating Adaptive Curricula through RMABs in Education Settings
by: Tio, Sidney, et al.
Published: (2024)
by: Tio, Sidney, et al.
Published: (2024)
Optimizing Ride-Pooling Operations with Extended Pickup and Drop-Off Flexibility
by: Jiang, Hao, et al.
Published: (2025)
by: Jiang, Hao, et al.
Published: (2025)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
Offline Safe Reinforcement Learning Using Trajectory Classification
by: Gong, Ze, et al.
Published: (2024)
by: Gong, Ze, et al.
Published: (2024)
On Minimizing Adversarial Counterfactual Error in Adversarial RL
by: Belaire, Roman, et al.
Published: (2024)
by: Belaire, Roman, et al.
Published: (2024)
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
by: Lu, Yuxiao, et al.
Published: (2024)
by: Lu, Yuxiao, et al.
Published: (2024)
Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement Learning
by: Hoang, Huy, et al.
Published: (2023)
by: Hoang, Huy, et al.
Published: (2023)
Automatic LLM Red Teaming
by: Belaire, Roman, et al.
Published: (2025)
by: Belaire, Roman, et al.
Published: (2025)
Imitating Cost-Constrained Behaviors in Reinforcement Learning
by: Shao, Qian, et al.
Published: (2024)
by: Shao, Qian, et al.
Published: (2024)
A Factored MDP Approach To Moving Target Defense With Dynamic Threat Modeling and Cost Efficiency
by: Bose, Megha, et al.
Published: (2024)
by: Bose, Megha, et al.
Published: (2024)
Principal-Agent Reward Shaping in MDPs
by: Ben-Porat, Omer, et al.
Published: (2023)
by: Ben-Porat, Omer, et al.
Published: (2023)
Towards Neural Network based Cognitive Models of Dynamic Decision-Making by Humans
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression
by: Ge, Zichang, et al.
Published: (2025)
by: Ge, Zichang, et al.
Published: (2025)
Regret-Based Defense in Adversarial Reinforcement Learning
by: Belaire, Roman, et al.
Published: (2023)
by: Belaire, Roman, et al.
Published: (2023)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2025)
by: Hoang, Huy, et al.
Published: (2025)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
by: Merrill, Scott, et al.
Published: (2026)
by: Merrill, Scott, et al.
Published: (2026)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Solving Long-run Average Reward Robust MDPs via Stochastic Games
by: Chatterjee, Krishnendu, et al.
Published: (2023)
by: Chatterjee, Krishnendu, et al.
Published: (2023)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
by: Jain, Gauri, et al.
Published: (2024)
by: Jain, Gauri, et al.
Published: (2024)
RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis
by: Wang, Jianwei, et al.
Published: (2025)
by: Wang, Jianwei, et al.
Published: (2025)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026)
by: Muni, Aneri, et al.
Published: (2026)
Privacy-Preserving AI-Enabled Decentralized Learning and Employment Records System
by: Xu, Yuqiao, et al.
Published: (2026)
by: Xu, Yuqiao, et al.
Published: (2026)
Decentralized AI-driven IoT Architecture for Privacy-Preserving and Latency-Optimized Healthcare in Pandemic and Critical Care Scenarios
by: Sammangi, Harsha, et al.
Published: (2025)
by: Sammangi, Harsha, et al.
Published: (2025)
Privacy-Preserving LLMs Routing
by: Wu, Xidong, et al.
Published: (2026)
by: Wu, Xidong, et al.
Published: (2026)
Declarative Privacy-Preserving Inference Queries
by: Guan, Hong, et al.
Published: (2024)
by: Guan, Hong, et al.
Published: (2024)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
by: Pukdee, Rattana, et al.
Published: (2026)
by: Pukdee, Rattana, et al.
Published: (2026)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
PAC Privacy Preserving Diffusion Models
by: Xu, Qipan, et al.
Published: (2023)
by: Xu, Qipan, et al.
Published: (2023)
Similar Items
-
Safety through feedback in Constrained RL
by: Chirra, Shashank Reddy, et al.
Published: (2024) -
On Discovering Algorithms for Adversarial Imitation Learning
by: Chirra, Shashank Reddy, et al.
Published: (2025) -
Efficient Unsupervised Environment Design through Hierarchical Policy Representation Learning
by: Li, Dexun, et al.
Published: (2026) -
PyVRP$^+$: LLM-Driven Metacognitive Heuristic Evolution for Hybrid Genetic Search in Vehicle Routing Problems
by: Malik, Manuj, et al.
Published: (2026) -
Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
by: Lu, Yuxiao, et al.
Published: (2023)