Principal-Agent Reward Shaping in MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Ben-Porat, Omer, Mansour, Yishay, Moshkovitz, Michal, Taitler, Boaz |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborating with GenAI: Incentives and Replacements
by: Taitler, Boaz, et al.
Published: (2025)
by: Taitler, Boaz, et al.
Published: (2025)
Selective Response Strategies for GenAI
by: Taitler, Boaz, et al.
Published: (2025)
by: Taitler, Boaz, et al.
Published: (2025)
Contests with Spillovers: Incentivizing Content Creation with GenAI
by: Ohayon, Sagi, et al.
Published: (2026)
by: Ohayon, Sagi, et al.
Published: (2026)
Data Sharing with a Generative AI Competitor
by: Taitler, Boaz, et al.
Published: (2025)
by: Taitler, Boaz, et al.
Published: (2025)
Braess's Paradox of Generative AI
by: Taitler, Boaz, et al.
Published: (2024)
by: Taitler, Boaz, et al.
Published: (2024)
Churn-Aware Recommendation Planning under Aggregated Preference Feedback
by: Keinan, Gur, et al.
Published: (2025)
by: Keinan, Gur, et al.
Published: (2025)
Modeling Churn in Recommender Systems with Aggregated Preferences
by: Keinan, Gur, et al.
Published: (2025)
by: Keinan, Gur, et al.
Published: (2025)
A Sequential Decision-Making Model for Perimeter Identification
by: Taitler, Ayal
Published: (2024)
by: Taitler, Ayal
Published: (2024)
Bandits with Single-Peaked Preferences and Limited Resources
by: Ben-Porat, Omer, et al.
Published: (2025)
by: Ben-Porat, Omer, et al.
Published: (2025)
Envious Explore and Exploit
by: Ben-Porat, Omer, et al.
Published: (2025)
by: Ben-Porat, Omer, et al.
Published: (2025)
From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums
by: Fono, Niv, et al.
Published: (2026)
by: Fono, Niv, et al.
Published: (2026)
Model-Driven Policy Optimization in Differentiable Simulators via Stochastic Exploration
by: Aroosh, Yuval, et al.
Published: (2026)
by: Aroosh, Yuval, et al.
Published: (2026)
Preserving the Privacy of Reward Functions in MDPs through Deception
by: Chirra, Shashank Reddy, et al.
Published: (2024)
by: Chirra, Shashank Reddy, et al.
Published: (2024)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Cooperative Reward Shaping for Multi-Agent Pathfinding
by: Song, Zhenyu, et al.
Published: (2024)
by: Song, Zhenyu, et al.
Published: (2024)
Solving Long-run Average Reward Robust MDPs via Stochastic Games
by: Chatterjee, Krishnendu, et al.
Published: (2023)
by: Chatterjee, Krishnendu, et al.
Published: (2023)
The Real Price of Bandit Information in Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
Fast Rates for Bandit PAC Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
From Kinematics to Dynamics: Learning to Refine Hybrid Plans for Physically Feasible Execution
by: Erez, Lidor, et al.
Published: (2026)
by: Erez, Lidor, et al.
Published: (2026)
Beyond Data Scarcity: A Frequency-Driven Framework for Zero-Shot Forecasting
by: Nochumsohn, Liran, et al.
Published: (2024)
by: Nochumsohn, Liran, et al.
Published: (2024)
Modeling Attrition in Recommender Systems with Departing Bandits
by: Ben-Porat, Omer, et al.
Published: (2022)
by: Ben-Porat, Omer, et al.
Published: (2022)
ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning
by: Abboud, Elie, et al.
Published: (2026)
by: Abboud, Elie, et al.
Published: (2026)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
Rising Rested MAB with Linear Drift
by: Amichay, Omer, et al.
Published: (2025)
by: Amichay, Omer, et al.
Published: (2025)
Agent-Specific Effects: A Causal Effect Propagation Analysis in Multi-Agent MDPs
by: Triantafyllou, Stelios, et al.
Published: (2023)
by: Triantafyllou, Stelios, et al.
Published: (2023)
Gradient-Free Training of Quantized Neural Networks
by: Cohen, Noa, et al.
Published: (2024)
by: Cohen, Noa, et al.
Published: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026)
by: Muni, Aneri, et al.
Published: (2026)
Regret Minimization and Convergence to Equilibria in General-sum Markov Games
by: Erez, Liad, et al.
Published: (2022)
by: Erez, Liad, et al.
Published: (2022)
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games
by: Afonso, António, et al.
Published: (2025)
by: Afonso, António, et al.
Published: (2025)
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
by: Barnea, Idan, et al.
Published: (2026)
by: Barnea, Idan, et al.
Published: (2026)
A Theory of Interpretable Approximations
by: Bressan, Marco, et al.
Published: (2024)
by: Bressan, Marco, et al.
Published: (2024)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
by: Wu, Boyuan
Published: (2025)
by: Wu, Boyuan
Published: (2025)
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Eluder-based Regret for Stochastic Contextual MDPs
by: Levy, Orin, et al.
Published: (2022)
by: Levy, Orin, et al.
Published: (2022)
Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory
by: Orogat, Abdelghny, et al.
Published: (2026)
by: Orogat, Abdelghny, et al.
Published: (2026)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
by: Erez, Liad, et al.
Published: (2026)
by: Erez, Liad, et al.
Published: (2026)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
by: Holmes, Ian, et al.
Published: (2025)
by: Holmes, Ian, et al.
Published: (2025)
Similar Items
-
Collaborating with GenAI: Incentives and Replacements
by: Taitler, Boaz, et al.
Published: (2025) -
Selective Response Strategies for GenAI
by: Taitler, Boaz, et al.
Published: (2025) -
Contests with Spillovers: Incentivizing Content Creation with GenAI
by: Ohayon, Sagi, et al.
Published: (2026) -
Data Sharing with a Generative AI Competitor
by: Taitler, Boaz, et al.
Published: (2025) -
Braess's Paradox of Generative AI
by: Taitler, Boaz, et al.
Published: (2024)