On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Williams, Marcus, Carroll, Micah, Narang, Adhyyan, Weisser, Constantin, Murphy, Brendan, Dragan, Anca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AI Alignment with Changing and Influenceable Reward Functions
von: Carroll, Micah, et al.
Veröffentlicht: (2024)
von: Carroll, Micah, et al.
Veröffentlicht: (2024)
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
von: Narang, Adhyyan, et al.
Veröffentlicht: (2024)
von: Narang, Adhyyan, et al.
Veröffentlicht: (2024)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
von: Howe, Nikolaus, et al.
Veröffentlicht: (2025)
von: Howe, Nikolaus, et al.
Veröffentlicht: (2025)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2024)
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2024)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)
von: Lang, Leon, et al.
Veröffentlicht: (2024)
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2023)
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2023)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Adversaries Can Misuse Combinations of Safe Models
von: Jones, Erik, et al.
Veröffentlicht: (2024)
von: Jones, Erik, et al.
Veröffentlicht: (2024)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
CTRL-Rec: Controlling Recommender Systems With Natural Language
von: Carroll, Micah, et al.
Veröffentlicht: (2025)
von: Carroll, Micah, et al.
Veröffentlicht: (2025)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2026)
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2026)
Not All News Is Equal: Topic- and Event-Conditional Sentiment from Finetuned LLMs for Aluminum Price Forecasting
von: Amorin, Alvaro Paredes, et al.
Veröffentlicht: (2026)
von: Amorin, Alvaro Paredes, et al.
Veröffentlicht: (2026)
Building Better Deception Probes Using Targeted Instruction Pairs
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
Training LLM Agents to Empower Humans
von: Ellis, Evan, et al.
Veröffentlicht: (2025)
von: Ellis, Evan, et al.
Veröffentlicht: (2025)
Aligning Robot and Human Representations
von: Bobu, Andreea, et al.
Veröffentlicht: (2023)
von: Bobu, Andreea, et al.
Veröffentlicht: (2023)
Learning from Streaming Data when Users Choose
von: Su, Jinyan, et al.
Veröffentlicht: (2024)
von: Su, Jinyan, et al.
Veröffentlicht: (2024)
Scaling Trends for Data Poisoning in LLMs
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
von: Narang, Adhyyan, et al.
Veröffentlicht: (2026)
von: Narang, Adhyyan, et al.
Veröffentlicht: (2026)
A Generalized Acquisition Function for Preference-based Reward Learning
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
Online SuBmodular + SuPermodular (BP) Maximization with Bandit Feedback
von: Narang, Adhyyan, et al.
Veröffentlicht: (2022)
von: Narang, Adhyyan, et al.
Veröffentlicht: (2022)
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
von: Wu, Jiaxing, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxing, et al.
Veröffentlicht: (2024)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
von: Pan, Michelle, et al.
Veröffentlicht: (2024)
von: Pan, Michelle, et al.
Veröffentlicht: (2024)
DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation
von: Mandi, Zhao, et al.
Veröffentlicht: (2025)
von: Mandi, Zhao, et al.
Veröffentlicht: (2025)
AssistanceZero: Scalably Solving Assistance Games
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2025)
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2025)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
von: Cheng, Ching-An, et al.
Veröffentlicht: (2024)
von: Cheng, Ching-An, et al.
Veröffentlicht: (2024)
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
von: Zolfaghari, Vahideh
Veröffentlicht: (2026)
von: Zolfaghari, Vahideh
Veröffentlicht: (2026)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
von: Hogan, Brendan R., et al.
Veröffentlicht: (2026)
von: Hogan, Brendan R., et al.
Veröffentlicht: (2026)
Deceptive Exploration in Multi-armed Bandits
von: Vurankaya, I. Arda, et al.
Veröffentlicht: (2025)
von: Vurankaya, I. Arda, et al.
Veröffentlicht: (2025)
Towards User-level Private Reinforcement Learning with Human Feedback
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
Learning to Model the World with Language
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
Prompt Optimization with Human Feedback
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
Transferable and Forecastable User Targeting Foundation Model
von: Dou, Bin, et al.
Veröffentlicht: (2024)
von: Dou, Bin, et al.
Veröffentlicht: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Exploitation Is All You Need... for Exploration
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
RL + Transformer = A General-Purpose Problem Solver
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
Learning to Assist Humans without Inferring Rewards
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
Information-theoretic Distinctions Between Deception and Confusion
von: Young, Robin
Veröffentlicht: (2025)
von: Young, Robin
Veröffentlicht: (2025)
HAEPO: History-Aggregated Exploratory Policy Optimization
von: Trivedi, Gaurish, et al.
Veröffentlicht: (2025)
von: Trivedi, Gaurish, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AI Alignment with Changing and Influenceable Reward Functions
von: Carroll, Micah, et al.
Veröffentlicht: (2024) -
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
von: Narang, Adhyyan, et al.
Veröffentlicht: (2024) -
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
von: Howe, Nikolaus, et al.
Veröffentlicht: (2025) -
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2024) -
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)