Towards Understanding Sycophancy in Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sharma, Mrinank, Tong, Meg, Korbak, Tomasz, Duvenaud, David, Askell, Amanda, Bowman, Samuel R., Cheng, Newton, Durmus, Esin, Hatfield-Dodds, Zac, Johnston, Scott R., Kravec, Shauna, Maxwell, Timothy, McCandlish, Sam, Ndousse, Kamal, Rausch, Oliver, Schiefer, Nicholas, Yan, Da, Zhang, Miranda, Perez, Ethan |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
par: Denison, Carson, et autres
Publié: (2024)
par: Denison, Carson, et autres
Publié: (2024)
Towards Measuring the Representation of Subjective Global Opinions in Language Models
par: Durmus, Esin, et autres
Publié: (2023)
par: Durmus, Esin, et autres
Publié: (2023)
Sabotage Evaluations for Frontier Models
par: Benton, Joe, et autres
Publié: (2024)
par: Benton, Joe, et autres
Publié: (2024)
Density estimation for ordinal biological sequences and its applications
par: Chen, Wei-Chia, et autres
Publié: (2024)
par: Chen, Wei-Chia, et autres
Publié: (2024)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
par: Hubinger, Evan, et autres
Publié: (2024)
par: Hubinger, Evan, et autres
Publié: (2024)
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
par: Maaz, Muhammad, et autres
Publié: (2025)
par: Maaz, Muhammad, et autres
Publié: (2025)
Aligning language models with human preferences
par: Korbak, Tomasz
Publié: (2024)
par: Korbak, Tomasz
Publié: (2024)
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
par: Sharma, Mrinank, et autres
Publié: (2026)
par: Sharma, Mrinank, et autres
Publié: (2026)
On learning functions over biological sequence space: relating Gaussian process priors, regularization, and gauge fixing
par: Petti, Samantha, et autres
Publié: (2025)
par: Petti, Samantha, et autres
Publié: (2025)
Improving Code Generation by Training with Natural Language Feedback
par: Chen, Angelica, et autres
Publié: (2023)
par: Chen, Angelica, et autres
Publié: (2023)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
par: Berglund, Lukas, et autres
Publié: (2023)
par: Berglund, Lukas, et autres
Publié: (2023)
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
par: Kuciński, Łukasz, et autres
Publié: (2021)
par: Kuciński, Łukasz, et autres
Publié: (2021)
The effects on structure of a momentum coupling between dark matter and quintessence
par: Candlish, G. N., et autres
Publié: (2025)
par: Candlish, G. N., et autres
Publié: (2025)
Forecasting Rare Language Model Behaviors
par: Jones, Erik, et autres
Publié: (2025)
par: Jones, Erik, et autres
Publié: (2025)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
par: Peng, Alwin, et autres
Publié: (2024)
par: Peng, Alwin, et autres
Publié: (2024)
Intergenerational Care in Local, Long‐Distance, and Transnational Families: The Role of Geographical Distance and Cross‐Border Separation on Subjective Care Burden
par: David Schiefer, et autres
Publié: (2024)
par: David Schiefer, et autres
Publié: (2024)
Training Language Models with Language Feedback at Scale
par: Scheurer, Jérémy, et autres
Publié: (2023)
par: Scheurer, Jérémy, et autres
Publié: (2023)
Compositional preference models for aligning LMs
par: Go, Dongyoung, et autres
Publié: (2023)
par: Go, Dongyoung, et autres
Publié: (2023)
NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
par: Gligoric, Kristina, et autres
Publié: (2024)
par: Gligoric, Kristina, et autres
Publié: (2024)
Stress-Testing Model Specs Reveals Character Differences among Language Models
par: Zhang, Jifan, et autres
Publié: (2025)
par: Zhang, Jifan, et autres
Publié: (2025)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
par: Natan, Shahar Ben, et autres
Publié: (2026)
par: Natan, Shahar Ben, et autres
Publié: (2026)
dempz/WAACHShelp: v1.4.2
par: Zac Dempsey
Publié: (2025)
par: Zac Dempsey
Publié: (2025)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
par: Youstra, Jack, et autres
Publié: (2025)
par: Youstra, Jack, et autres
Publié: (2025)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
par: Kasneci, Enkelejda, et autres
Publié: (2026)
par: Kasneci, Enkelejda, et autres
Publié: (2026)
Abundance and Hysteresis Data Accompanying Magnetic Properties of Marine Sediments: Bulk Samples versus Grain Size Separates
par: Hatfield, Robert
Publié: (2025)
par: Hatfield, Robert
Publié: (2025)
Abundance and Hysteresis Data Accompanying Magnetic Properties of Marine Sediments: Bulk Samples versus Grain Size Separates
par: Hatfield, Robert
Publié: (2025)
par: Hatfield, Robert
Publié: (2025)
Partnerships in Information Services: The Contract Library.
par: Hatfield, Deborah
Publié: (1994)
par: Hatfield, Deborah
Publié: (1994)
Balint Groups: A Catalytic Collaboration Between General Practice and Psychoanalysis
par: Robert V. Dyer, et autres
Publié: (2026)
par: Robert V. Dyer, et autres
Publié: (2026)
Feelings of Guilt When Caring for Parents Across Borders: The Role of Gender and Country‐Specific Care Systems and Norms
par: David Schiefer, et autres
Publié: (2025)
par: David Schiefer, et autres
Publié: (2025)
Nahrungsergänzungsmittel unter der Lupe
par: H. Osiander‐Fuchs, et autres
Publié: (2024)
par: H. Osiander‐Fuchs, et autres
Publié: (2024)
Adaptive Physics-informed Neural Networks: A Survey
par: Torres, Edgar, et autres
Publié: (2025)
par: Torres, Edgar, et autres
Publié: (2025)
Evening Services of Junior Colleges.
par: Hatfield, Thomas M., et autres
Publié: (1970)
par: Hatfield, Thomas M., et autres
Publié: (1970)
How RLHF Amplifies Sycophancy
par: Shapira, Itai, et autres
Publié: (2026)
par: Shapira, Itai, et autres
Publié: (2026)
Correção de transplante capilar inestético
par: Renata Indelicato Zac
Publié: (2011)
par: Renata Indelicato Zac
Publié: (2011)
Diphyllobothriid cestodes form the Hawaiian monk seal, Monachus schauinslandi Matschie, from Midway Atoll
par: Rausch, R. L
Publié: (1969)
par: Rausch, R. L
Publié: (1969)
Music and the Cultural Production of Scale
par: Dodds, Phil
Publié: (2023)
par: Dodds, Phil
Publié: (2023)
Curricular Injustice: How U.S. Medical Schools Reproduce Inequalities. By L.Olsen, New York: Columbia University Press, 2024. 312 pp. $140 (hard); $35 (pbk); $34.99 (ebk). ISBN: 978‐0‐23‐120787‐4
par: Catherine Dodds
Publié: (2026)
par: Catherine Dodds
Publié: (2026)
St. Kitts at a Crossroad
par: Dodds, Rachel
Publié: (2008)
par: Dodds, Rachel
Publié: (2008)
St. Kitts at a Crossroad / Rachel Dodds, Jerome L. McElroy
par: Dodds, Rachel
Publié: (2005)
par: Dodds, Rachel
Publié: (2005)
Evaluando la relación entre el consumo televisivo y las actitudes hacia personas viviendo con VIH/sida en Chile
par: Tomás Dodds
Publié: (2017)
par: Tomás Dodds
Publié: (2017)
Documents similaires
-
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
par: Denison, Carson, et autres
Publié: (2024) -
Towards Measuring the Representation of Subjective Global Opinions in Language Models
par: Durmus, Esin, et autres
Publié: (2023) -
Sabotage Evaluations for Frontier Models
par: Benton, Joe, et autres
Publié: (2024) -
Density estimation for ordinal biological sequences and its applications
par: Chen, Wei-Chia, et autres
Publié: (2024) -
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
par: Hubinger, Evan, et autres
Publié: (2024)