From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Wei, Huang, Zhen, Xie, Liang, Lin, Binbin, Li, Houqiang, Lu, Le, Tian, Xinmei, Cai, Deng, Zhang, Yonggang, Wang, Wenxiao, Shen, Xu, Ye, Jieping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
by: Dai, Rui, et al.
Published: (2025)
by: Dai, Rui, et al.
Published: (2025)
Interpreting and Improving Large Language Models in Arithmetic Calculation
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Delving into the Reversal Curse: How Far Can Large Language Models Generalize?
by: Lin, Zhengkai, et al.
Published: (2024)
by: Lin, Zhengkai, et al.
Published: (2024)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025)
by: Yan, Shaotian, et al.
Published: (2025)
Concise and Organized Perception Facilitates Reasoning in Large Language Models
by: Liu, Junjie, et al.
Published: (2023)
by: Liu, Junjie, et al.
Published: (2023)
GeoCAD: Local Geometry-Controllable CAD Generation with Large Language Models
by: Zhang, Zhanwei, et al.
Published: (2025)
by: Zhang, Zhanwei, et al.
Published: (2025)
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
by: Xu, Shixiong, et al.
Published: (2025)
by: Xu, Shixiong, et al.
Published: (2025)
SciPIP: An LLM-based Scientific Paper Idea Proposer
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
by: Huang, Chenxi, et al.
Published: (2025)
by: Huang, Chenxi, et al.
Published: (2025)
When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
Controlling Thinking Speed in Reasoning Models
by: Lin, Zhengkai, et al.
Published: (2025)
by: Lin, Zhengkai, et al.
Published: (2025)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Career Aspirations of MLS Students; Yes, the Women Are as Ambitious as the Men.
by: Harris, Roma M.
Published: (1986)
by: Harris, Roma M.
Published: (1986)
Digital Yes‐Men: How to Deal With Sycophantic Military AI ?
by: Jonathan Kwik
Published: (2025)
by: Jonathan Kwik
Published: (2025)
Enhancing Spatial Reasoning through Visual and Textual Thinking
by: Liang, Xun, et al.
Published: (2025)
by: Liang, Xun, et al.
Published: (2025)
Instance-adaptive Zero-shot Chain-of-Thought Prompting
by: Yuan, Xiaosong, et al.
Published: (2024)
by: Yuan, Xiaosong, et al.
Published: (2024)
Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
by: Natan, Shahar Ben, et al.
Published: (2026)
by: Natan, Shahar Ben, et al.
Published: (2026)
Should the Internet Be Censored? Yes! Yes! Yes! Yes!
by: Taylor, Bruce
Published: (1998)
by: Taylor, Bruce
Published: (1998)
Enhancing Target-unspecific Tasks through a Features Matrix
by: Cui, Fangming, et al.
Published: (2025)
by: Cui, Fangming, et al.
Published: (2025)
Towards Quantum Accelerated Large-scale Topology Optimization
by: Ye, Zisheng, et al.
Published: (2025)
by: Ye, Zisheng, et al.
Published: (2025)
OBMO: One Bounding Box Multiple Objects for Monocular 3D Object Detection
by: Huang, Chenxi, et al.
Published: (2022)
by: Huang, Chenxi, et al.
Published: (2022)
Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow
by: Zhang, Chengsheng, et al.
Published: (2026)
by: Zhang, Chengsheng, et al.
Published: (2026)
AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization
by: Xu, Shixiong, et al.
Published: (2024)
by: Xu, Shixiong, et al.
Published: (2024)
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
by: Fan, Sinan, et al.
Published: (2025)
by: Fan, Sinan, et al.
Published: (2025)
TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
by: Zhang, Yuxiang, et al.
Published: (2025)
by: Zhang, Yuxiang, et al.
Published: (2025)
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
by: Shah, Arya, et al.
Published: (2026)
by: Shah, Arya, et al.
Published: (2026)
PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
by: Çelebi, Yusuf, et al.
Published: (2025)
by: Çelebi, Yusuf, et al.
Published: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
by: Yuan, Botai, et al.
Published: (2025)
by: Yuan, Botai, et al.
Published: (2025)
Robust Training of Federated Models with Extremely Label Deficiency
by: Zhang, Yonggang, et al.
Published: (2024)
by: Zhang, Yonggang, et al.
Published: (2024)
Detecting Generated Images by Fitting Natural Image Distributions
by: Zhang, Yonggang, et al.
Published: (2025)
by: Zhang, Yonggang, et al.
Published: (2025)
G2LTraj: A Global-to-Local Generation Approach for Trajectory Prediction
by: Zhang, Zhanwei, et al.
Published: (2024)
by: Zhang, Zhanwei, et al.
Published: (2024)
Sycophancy in Large Language Models: Causes and Mitigations
by: Malmqvist, Lars
Published: (2024)
by: Malmqvist, Lars
Published: (2024)
Do Women Legislators Legislate Differently Than Men on Gun‐Related Policy? A Suggestive Yes
by: Patrick Cunha Silva, et al.
Published: (2026)
by: Patrick Cunha Silva, et al.
Published: (2026)
MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems
by: Cai, Qianshu, et al.
Published: (2026)
by: Cai, Qianshu, et al.
Published: (2026)
Pinpointing Plastic Water Lines
by: Marc Bracken, et al.
Published: (2026)
by: Marc Bracken, et al.
Published: (2026)
Pinpointing when and how of teledermatology
by: Paola Pasquali, et al.
Published: (2025)
by: Paola Pasquali, et al.
Published: (2025)
Similar Items
-
Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
by: Dai, Rui, et al.
Published: (2025) -
Interpreting and Improving Large Language Models in Arithmetic Calculation
by: Zhang, Wei, et al.
Published: (2024) -
Delving into the Reversal Curse: How Far Can Large Language Models Generalize?
by: Lin, Zhengkai, et al.
Published: (2024) -
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
by: Xiao, Yuxin, et al.
Published: (2024) -
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025)