On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
Fuente:
arXiv
Salvato in:
| Autori principali: | Williams, Marcus, Carroll, Micah, Narang, Adhyyan, Weisser, Constantin, Murphy, Brendan, Dragan, Anca |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AI Alignment with Changing and Influenceable Reward Functions
di: Carroll, Micah, et al.
Pubblicazione: (2024)
di: Carroll, Micah, et al.
Pubblicazione: (2024)
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
di: Narang, Adhyyan, et al.
Pubblicazione: (2024)
di: Narang, Adhyyan, et al.
Pubblicazione: (2024)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
di: Howe, Nikolaus, et al.
Pubblicazione: (2025)
di: Howe, Nikolaus, et al.
Pubblicazione: (2025)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024)
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
di: Lang, Leon, et al.
Pubblicazione: (2024)
di: Lang, Leon, et al.
Pubblicazione: (2024)
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2023)
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2023)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
di: Hong, Joey, et al.
Pubblicazione: (2024)
di: Hong, Joey, et al.
Pubblicazione: (2024)
Adversaries Can Misuse Combinations of Safe Models
di: Jones, Erik, et al.
Pubblicazione: (2024)
di: Jones, Erik, et al.
Pubblicazione: (2024)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
di: Myers, Vivek, et al.
Pubblicazione: (2024)
di: Myers, Vivek, et al.
Pubblicazione: (2024)
CTRL-Rec: Controlling Recommender Systems With Natural Language
di: Carroll, Micah, et al.
Pubblicazione: (2025)
di: Carroll, Micah, et al.
Pubblicazione: (2025)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
di: Hong, Joey, et al.
Pubblicazione: (2024)
di: Hong, Joey, et al.
Pubblicazione: (2024)
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
di: Qiu, Tianyi Alex, et al.
Pubblicazione: (2026)
di: Qiu, Tianyi Alex, et al.
Pubblicazione: (2026)
Not All News Is Equal: Topic- and Event-Conditional Sentiment from Finetuned LLMs for Aluminum Price Forecasting
di: Amorin, Alvaro Paredes, et al.
Pubblicazione: (2026)
di: Amorin, Alvaro Paredes, et al.
Pubblicazione: (2026)
Building Better Deception Probes Using Targeted Instruction Pairs
di: Natarajan, Vikram, et al.
Pubblicazione: (2026)
di: Natarajan, Vikram, et al.
Pubblicazione: (2026)
Training LLM Agents to Empower Humans
di: Ellis, Evan, et al.
Pubblicazione: (2025)
di: Ellis, Evan, et al.
Pubblicazione: (2025)
Aligning Robot and Human Representations
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
Learning from Streaming Data when Users Choose
di: Su, Jinyan, et al.
Pubblicazione: (2024)
di: Su, Jinyan, et al.
Pubblicazione: (2024)
Scaling Trends for Data Poisoning in LLMs
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing
di: Narang, Adhyyan, et al.
Pubblicazione: (2026)
di: Narang, Adhyyan, et al.
Pubblicazione: (2026)
A Generalized Acquisition Function for Preference-based Reward Learning
di: Ellis, Evan, et al.
Pubblicazione: (2024)
di: Ellis, Evan, et al.
Pubblicazione: (2024)
Online SuBmodular + SuPermodular (BP) Maximization with Bandit Feedback
di: Narang, Adhyyan, et al.
Pubblicazione: (2022)
di: Narang, Adhyyan, et al.
Pubblicazione: (2022)
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
di: Wu, Jiaxing, et al.
Pubblicazione: (2024)
di: Wu, Jiaxing, et al.
Pubblicazione: (2024)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
di: Pan, Michelle, et al.
Pubblicazione: (2024)
di: Pan, Michelle, et al.
Pubblicazione: (2024)
DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation
di: Mandi, Zhao, et al.
Pubblicazione: (2025)
di: Mandi, Zhao, et al.
Pubblicazione: (2025)
AssistanceZero: Scalably Solving Assistance Games
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2025)
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2025)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
di: Zolfaghari, Vahideh
Pubblicazione: (2026)
di: Zolfaghari, Vahideh
Pubblicazione: (2026)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
di: Hogan, Brendan R., et al.
Pubblicazione: (2026)
di: Hogan, Brendan R., et al.
Pubblicazione: (2026)
Deceptive Exploration in Multi-armed Bandits
di: Vurankaya, I. Arda, et al.
Pubblicazione: (2025)
di: Vurankaya, I. Arda, et al.
Pubblicazione: (2025)
Towards User-level Private Reinforcement Learning with Human Feedback
di: Zhang, Jiaming, et al.
Pubblicazione: (2025)
di: Zhang, Jiaming, et al.
Pubblicazione: (2025)
Learning to Model the World with Language
di: Lin, Jessy, et al.
Pubblicazione: (2023)
di: Lin, Jessy, et al.
Pubblicazione: (2023)
Prompt Optimization with Human Feedback
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2024)
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2024)
Transferable and Forecastable User Targeting Foundation Model
di: Dou, Bin, et al.
Pubblicazione: (2024)
di: Dou, Bin, et al.
Pubblicazione: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
Exploitation Is All You Need... for Exploration
di: Rentschler, Micah, et al.
Pubblicazione: (2025)
di: Rentschler, Micah, et al.
Pubblicazione: (2025)
RL + Transformer = A General-Purpose Problem Solver
di: Rentschler, Micah, et al.
Pubblicazione: (2025)
di: Rentschler, Micah, et al.
Pubblicazione: (2025)
Learning to Assist Humans without Inferring Rewards
di: Myers, Vivek, et al.
Pubblicazione: (2024)
di: Myers, Vivek, et al.
Pubblicazione: (2024)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
di: Zhang, Eva, et al.
Pubblicazione: (2024)
di: Zhang, Eva, et al.
Pubblicazione: (2024)
Information-theoretic Distinctions Between Deception and Confusion
di: Young, Robin
Pubblicazione: (2025)
di: Young, Robin
Pubblicazione: (2025)
HAEPO: History-Aggregated Exploratory Policy Optimization
di: Trivedi, Gaurish, et al.
Pubblicazione: (2025)
di: Trivedi, Gaurish, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AI Alignment with Changing and Influenceable Reward Functions
di: Carroll, Micah, et al.
Pubblicazione: (2024) -
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning
di: Narang, Adhyyan, et al.
Pubblicazione: (2024) -
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
di: Howe, Nikolaus, et al.
Pubblicazione: (2025) -
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024) -
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
di: Lang, Leon, et al.
Pubblicazione: (2024)