Linear Probe Penalties Reduce LLM Sycophancy
Fuente:
arXiv
Salvato in:
| Autori principali: | Papadatos, Henry, Freedman, Rachel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SycEval: Evaluating LLM Sycophancy
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
Sycophancy Hides Linearly in the Attention Heads
di: Genadi, Rifo, et al.
Pubblicazione: (2026)
di: Genadi, Rifo, et al.
Pubblicazione: (2026)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
di: Kasneci, Enkelejda, et al.
Pubblicazione: (2026)
di: Kasneci, Enkelejda, et al.
Pubblicazione: (2026)
Mapping AI Benchmark Data to Quantitative Risk Estimates Through Expert Elicitation
di: Murray, Malcolm, et al.
Pubblicazione: (2025)
di: Murray, Malcolm, et al.
Pubblicazione: (2025)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
di: Chang, Edward Y.
Pubblicazione: (2026)
di: Chang, Edward Y.
Pubblicazione: (2026)
How RLHF Amplifies Sycophancy
di: Shapira, Itai, et al.
Pubblicazione: (2026)
di: Shapira, Itai, et al.
Pubblicazione: (2026)
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management
di: Campos, Simeon, et al.
Pubblicazione: (2025)
di: Campos, Simeon, et al.
Pubblicazione: (2025)
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
di: Zhao, Zhenyu, et al.
Pubblicazione: (2026)
di: Zhao, Zhenyu, et al.
Pubblicazione: (2026)
Moral Sycophancy in Vision Language Models
di: Rabby, Shadman, et al.
Pubblicazione: (2026)
di: Rabby, Shadman, et al.
Pubblicazione: (2026)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
di: Natan, Shahar Ben, et al.
Pubblicazione: (2026)
di: Natan, Shahar Ben, et al.
Pubblicazione: (2026)
When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models
di: Li, Jiechen, et al.
Pubblicazione: (2026)
di: Li, Jiechen, et al.
Pubblicazione: (2026)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
di: Atwell, Katherine, et al.
Pubblicazione: (2025)
di: Atwell, Katherine, et al.
Pubblicazione: (2025)
Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
di: Törnberg, Petter, et al.
Pubblicazione: (2026)
di: Törnberg, Petter, et al.
Pubblicazione: (2026)
Sycophancy in Large Language Models: Causes and Mitigations
di: Malmqvist, Lars
Pubblicazione: (2024)
di: Malmqvist, Lars
Pubblicazione: (2024)
Consistency Training Helps Stop Sycophancy and Jailbreaks
di: Irpan, Alex, et al.
Pubblicazione: (2025)
di: Irpan, Alex, et al.
Pubblicazione: (2025)
The Silicon Mirror: Dynamic Behavioral Gating for Anti-Sycophancy in LLM Agents
di: Shah, Harshee Jignesh
Pubblicazione: (2026)
di: Shah, Harshee Jignesh
Pubblicazione: (2026)
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
di: Wang, Libo
Pubblicazione: (2024)
di: Wang, Libo
Pubblicazione: (2024)
PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models
di: Rahman, A. B. M. Ashikur, et al.
Pubblicazione: (2025)
di: Rahman, A. B. M. Ashikur, et al.
Pubblicazione: (2025)
Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition
di: Mohsin, Muhammad Ahmed, et al.
Pubblicazione: (2026)
di: Mohsin, Muhammad Ahmed, et al.
Pubblicazione: (2026)
Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
di: Beigi, Mohammad, et al.
Pubblicazione: (2025)
di: Beigi, Mohammad, et al.
Pubblicazione: (2025)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
Evaluating the Goal-Directedness of Large Language Models
di: Everitt, Tom, et al.
Pubblicazione: (2025)
di: Everitt, Tom, et al.
Pubblicazione: (2025)
Active teacher selection for reward learning
di: Freedman, Rachel, et al.
Pubblicazione: (2023)
di: Freedman, Rachel, et al.
Pubblicazione: (2023)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
di: Denison, Carson, et al.
Pubblicazione: (2024)
di: Denison, Carson, et al.
Pubblicazione: (2024)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct
di: Ye, Meryl, et al.
Pubblicazione: (2026)
di: Ye, Meryl, et al.
Pubblicazione: (2026)
Rhetorical Questions in LLM Representations: A Linear Probing Study
di: Yao, Louie Hong, et al.
Pubblicazione: (2026)
di: Yao, Louie Hong, et al.
Pubblicazione: (2026)
Accounting for Sycophancy in Language Model Uncertainty Estimation
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
di: Jain, Raunak
Pubblicazione: (2026)
di: Jain, Raunak
Pubblicazione: (2026)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
di: Xu, Juangui, et al.
Pubblicazione: (2025)
di: Xu, Juangui, et al.
Pubblicazione: (2025)
Generative AI in Managerial Decision-Making: Redefining Boundaries through Ambiguity Resolution and Sycophancy Analysis
di: Birim, Sule Ozturk, et al.
Pubblicazione: (2026)
di: Birim, Sule Ozturk, et al.
Pubblicazione: (2026)
Towards Understanding Sycophancy in Language Models
di: Sharma, Mrinank, et al.
Pubblicazione: (2023)
di: Sharma, Mrinank, et al.
Pubblicazione: (2023)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
di: Cencerrado, Iván Vicente Moreno, et al.
Pubblicazione: (2025)
di: Cencerrado, Iván Vicente Moreno, et al.
Pubblicazione: (2025)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
di: Pandey, Sanskar, et al.
Pubblicazione: (2025)
di: Pandey, Sanskar, et al.
Pubblicazione: (2025)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
di: Hong, Rui, et al.
Pubblicazione: (2026)
di: Hong, Rui, et al.
Pubblicazione: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
MambaKick: Early Penalty Direction Prediction from HAR Embeddings
di: Velesaca, Henry O., et al.
Pubblicazione: (2026)
di: Velesaca, Henry O., et al.
Pubblicazione: (2026)
Do Linear Probes Generalize Better in Persona Coordinates?
di: Mahadik, Prasad, et al.
Pubblicazione: (2026)
di: Mahadik, Prasad, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SycEval: Evaluating LLM Sycophancy
di: Fanous, Aaron, et al.
Pubblicazione: (2025) -
Sycophancy Hides Linearly in the Attention Heads
di: Genadi, Rifo, et al.
Pubblicazione: (2026) -
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
di: Kasneci, Enkelejda, et al.
Pubblicazione: (2026) -
Mapping AI Benchmark Data to Quantitative Risk Estimates Through Expert Elicitation
di: Murray, Malcolm, et al.
Pubblicazione: (2025) -
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
di: Chang, Edward Y.
Pubblicazione: (2026)