Language Models Learn to Mislead Humans via RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Jiaxin, Zhong, Ruiqi, Khan, Akbir, Perez, Ethan, Steinhardt, Jacob, Huang, Minlie, Bowman, Samuel R., He, He, Feng, Shi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Task Decomposition to Assist Humans in Competitive Programming
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Unsupervised Elicitation of Language Models
von: Wen, Jiaxin, et al.
Veröffentlicht: (2025)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2025)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
von: Jones, Erik, et al.
Veröffentlicht: (2025)
von: Jones, Erik, et al.
Veröffentlicht: (2025)
Factorio Learning Environment
von: Hopkins, Jack, et al.
Veröffentlicht: (2025)
von: Hopkins, Jack, et al.
Veröffentlicht: (2025)
Debating with More Persuasive LLMs Leads to More Truthful Answers
von: Khan, Akbir, et al.
Veröffentlicht: (2024)
von: Khan, Akbir, et al.
Veröffentlicht: (2024)
Spontaneous Reward Hacking in Iterative Self-Refinement
von: Pan, Jane, et al.
Veröffentlicht: (2024)
von: Pan, Jane, et al.
Veröffentlicht: (2024)
Approaching Human-Level Forecasting with Language Models
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
von: Leng, Jixuan, et al.
Veröffentlicht: (2024)
von: Leng, Jixuan, et al.
Veröffentlicht: (2024)
Mass-Producing Failures of Multimodal Systems with Language Models
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
What Artificial Neural Networks Can Tell Us About Human Language Acquisition
von: Warstadt, Alex, et al.
Veröffentlicht: (2022)
von: Warstadt, Alex, et al.
Veröffentlicht: (2022)
Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
von: Pfau, Jacob, et al.
Veröffentlicht: (2024)
von: Pfau, Jacob, et al.
Veröffentlicht: (2024)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
von: Feng, Andrew Zhuoer, et al.
Veröffentlicht: (2026)
Large Language Models as Misleading Assistants in Conversation
von: Hou, Betty Li, et al.
Veröffentlicht: (2024)
von: Hou, Betty Li, et al.
Veröffentlicht: (2024)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026)
von: Luo, Grace, et al.
Veröffentlicht: (2026)
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
von: Bihani, Geetanjali, et al.
Veröffentlicht: (2024)
von: Bihani, Geetanjali, et al.
Veröffentlicht: (2024)
Language Model Circuits Are Sparse in the Neuron Basis
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
von: Arora, Aryaman, et al.
Veröffentlicht: (2026)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
Towards Optimal Learning of Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
Language Model Decoding as Direct Metrics Optimization
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
LLM Evaluators Recognize and Favor Their Own Generations
von: Panickssery, Arjun, et al.
Veröffentlicht: (2024)
von: Panickssery, Arjun, et al.
Veröffentlicht: (2024)
Describing Differences in Image Sets with Natural Language
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
von: Yin, Yueqin, et al.
Veröffentlicht: (2025)
von: Yin, Yueqin, et al.
Veröffentlicht: (2025)
Data Selection via Optimal Control for Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models
von: Dunlap, Lisa, et al.
Veröffentlicht: (2024)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Task Decomposition to Assist Humans in Competitive Programming
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024) -
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024) -
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024) -
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023) -
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)