Policy Teaching via Data Poisoning in Learning from Human Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Nika, Andi, Nöther, Jonathan, Mandal, Debmalya, Kamalaruban, Parameswaran, Singla, Adish, Radanović, Goran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
by: Nika, Andi, et al.
Published: (2024)
by: Nika, Andi, et al.
Published: (2024)
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
by: Nika, Andi, et al.
Published: (2026)
by: Nika, Andi, et al.
Published: (2026)
Corruption Robust Offline Reinforcement Learning with Human Feedback
by: Mandal, Debmalya, et al.
Published: (2024)
by: Mandal, Debmalya, et al.
Published: (2024)
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
by: Nika, Andi, et al.
Published: (2024)
by: Nika, Andi, et al.
Published: (2024)
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
by: Nöther, Jonathan, et al.
Published: (2025)
by: Nöther, Jonathan, et al.
Published: (2025)
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
by: Nöther, Jonathan, et al.
Published: (2025)
by: Nöther, Jonathan, et al.
Published: (2025)
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
by: Nöther, Jonathan, et al.
Published: (2026)
by: Nöther, Jonathan, et al.
Published: (2026)
Informativeness of Reward Functions in Reinforcement Learning
by: Devidze, Rati, et al.
Published: (2024)
by: Devidze, Rati, et al.
Published: (2024)
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
by: Tzannetos, Georgios, et al.
Published: (2024)
by: Tzannetos, Georgios, et al.
Published: (2024)
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
by: Tzannetos, Georgios, et al.
Published: (2025)
by: Tzannetos, Georgios, et al.
Published: (2025)
Sparse Offline Reinforcement Learning with Corruption Robustness
by: Tran, Nam Phuong, et al.
Published: (2025)
by: Tran, Nam Phuong, et al.
Published: (2025)
Performative Reinforcement Learning with Linear Markov Decision Process
by: Mandal, Debmalya, et al.
Published: (2024)
by: Mandal, Debmalya, et al.
Published: (2024)
Distributionally Robust Reinforcement Learning with Human Feedback
by: Mandal, Debmalya, et al.
Published: (2025)
by: Mandal, Debmalya, et al.
Published: (2025)
Inference-Time Personalized Alignment with a Few User Preference Queries
by: Pădurean, Victor-Alexandru, et al.
Published: (2025)
by: Pădurean, Victor-Alexandru, et al.
Published: (2025)
On Corruption-Robustness in Performative Reinforcement Learning
by: Pollatos, Vasilis, et al.
Published: (2025)
by: Pollatos, Vasilis, et al.
Published: (2025)
Learning Embeddings for Sequential Tasks Using Population of Agents
by: Mahajan, Mridul, et al.
Published: (2023)
by: Mahajan, Mridul, et al.
Published: (2023)
Performative Reinforcement Learning in Gradually Shifting Environments
by: Rank, Ben, et al.
Published: (2024)
by: Rank, Ben, et al.
Published: (2024)
Stochastic Principal-Agent Problems: Efficient Computation and Learning
by: Gan, Jiarui, et al.
Published: (2023)
by: Gan, Jiarui, et al.
Published: (2023)
Independent Learning in Performative Markov Potential Games
by: Sahitaj, Rilind, et al.
Published: (2025)
by: Sahitaj, Rilind, et al.
Published: (2025)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
by: Sasnauskas, Paulius, et al.
Published: (2025)
by: Sasnauskas, Paulius, et al.
Published: (2025)
Reward Design for Justifiable Sequential Decision-Making
by: Sukovic, Aleksa, et al.
Published: (2024)
by: Sukovic, Aleksa, et al.
Published: (2024)
AgenticRed: Evolving Agentic Systems for Red-Teaming
by: Yuan, Jiayi, et al.
Published: (2026)
by: Yuan, Jiayi, et al.
Published: (2026)
Adversarially Robust Decision Transformer
by: Tang, Xiaohang, et al.
Published: (2024)
by: Tang, Xiaohang, et al.
Published: (2024)
Strategyproof Reinforcement Learning from Human Feedback
by: Buening, Thomas Kleine, et al.
Published: (2025)
by: Buening, Thomas Kleine, et al.
Published: (2025)
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
by: Kotalwar, Nachiket, et al.
Published: (2024)
by: Kotalwar, Nachiket, et al.
Published: (2024)
Learning Personalized Decision Support Policies
by: Bhatt, Umang, et al.
Published: (2023)
by: Bhatt, Umang, et al.
Published: (2023)
Formal Models of Active Learning from Contrastive Examples
by: Mansouri, Farnam, et al.
Published: (2025)
by: Mansouri, Farnam, et al.
Published: (2025)
Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human Input
by: Peng, Andi, et al.
Published: (2024)
by: Peng, Andi, et al.
Published: (2024)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
by: Radmehr, Bahar, et al.
Published: (2024)
by: Radmehr, Bahar, et al.
Published: (2024)
Contextual Combinatorial Bandits with Changing Action Sets via Gaussian Processes
by: Nika, Andi, et al.
Published: (2021)
by: Nika, Andi, et al.
Published: (2021)
Agent-Specific Effects: A Causal Effect Propagation Analysis in Multi-Agent MDPs
by: Triantafyllou, Stelios, et al.
Published: (2023)
by: Triantafyllou, Stelios, et al.
Published: (2023)
Learning Half-Spaces from Perturbed Contrastive Examples
by: Ravari, Aryan Alavi Razavi, et al.
Published: (2026)
by: Ravari, Aryan Alavi Razavi, et al.
Published: (2026)
Emergent Bias and Fairness in Multi-Agent Decision Systems
by: Madigan, Maeve, et al.
Published: (2025)
by: Madigan, Maeve, et al.
Published: (2025)
Evaluating Fairness in Transaction Fraud Models: Fairness Metrics, Bias Audits, and Challenges
by: Kamalaruban, Parameswaran, et al.
Published: (2024)
by: Kamalaruban, Parameswaran, et al.
Published: (2024)
Beyond Grids: Multi-objective Bayesian Optimization With Adaptive Discretization
by: Nika, Andi, et al.
Published: (2020)
by: Nika, Andi, et al.
Published: (2020)
Neural Task Synthesis for Visual Programming
by: Pădurean, Victor-Alexandru, et al.
Published: (2023)
by: Pădurean, Victor-Alexandru, et al.
Published: (2023)
Reinforcement Learning for Durable Algorithmic Recourse
by: Ceccon, Marina, et al.
Published: (2025)
by: Ceccon, Marina, et al.
Published: (2025)
Fairness-Aware Low-Rank Adaptation Under Demographic Privacy Constraints
by: Kamalaruban, Parameswaran, et al.
Published: (2025)
by: Kamalaruban, Parameswaran, et al.
Published: (2025)
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
by: Gölz, Paul, et al.
Published: (2025)
by: Gölz, Paul, et al.
Published: (2025)
COPR: Continual Human Preference Learning via Optimal Policy Regularization
by: Zhang, Han, et al.
Published: (2024)
by: Zhang, Han, et al.
Published: (2024)
Similar Items
-
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
by: Nika, Andi, et al.
Published: (2024) -
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
by: Nika, Andi, et al.
Published: (2026) -
Corruption Robust Offline Reinforcement Learning with Human Feedback
by: Mandal, Debmalya, et al.
Published: (2024) -
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
by: Nika, Andi, et al.
Published: (2024) -
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
by: Nöther, Jonathan, et al.
Published: (2025)