SycEval: Evaluating LLM Sycophancy
Fuente:
arXiv
Salvato in:
| Autori principali: | Fanous, Aaron, Goldberg, Jacob, Agarwal, Ank A., Lin, Joanna, Zhou, Anson, Daneshjou, Roxana, Koyejo, Sanmi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
Best Practices for Large Language Models in Radiology
di: Bluethgen, Christian, et al.
Pubblicazione: (2024)
di: Bluethgen, Christian, et al.
Pubblicazione: (2024)
Stop Automating Peer Review Without Rigorous Evaluation
di: Baumann, Joachim, et al.
Pubblicazione: (2026)
di: Baumann, Joachim, et al.
Pubblicazione: (2026)
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
A Framework for Objective-Driven Dynamical Stochastic Fields
di: Zhang, Yibo Jacky, et al.
Pubblicazione: (2025)
di: Zhang, Yibo Jacky, et al.
Pubblicazione: (2025)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
di: Salaudeen, Olawale, et al.
Pubblicazione: (2025)
di: Salaudeen, Olawale, et al.
Pubblicazione: (2025)
Reliable and Efficient Amortized Model-based Evaluation
di: Truong, Sang, et al.
Pubblicazione: (2025)
di: Truong, Sang, et al.
Pubblicazione: (2025)
SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care
di: Peng, Dongshen, et al.
Pubblicazione: (2026)
di: Peng, Dongshen, et al.
Pubblicazione: (2026)
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
di: Hariharan, Suhas, et al.
Pubblicazione: (2024)
di: Hariharan, Suhas, et al.
Pubblicazione: (2024)
Linear Probe Penalties Reduce LLM Sycophancy
di: Papadatos, Henry, et al.
Pubblicazione: (2024)
di: Papadatos, Henry, et al.
Pubblicazione: (2024)
BiasICL: In-Context Learning and Demographic Biases of Vision Language Models
di: Xu, Sonnet, et al.
Pubblicazione: (2025)
di: Xu, Sonnet, et al.
Pubblicazione: (2025)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
di: Kasneci, Enkelejda, et al.
Pubblicazione: (2026)
di: Kasneci, Enkelejda, et al.
Pubblicazione: (2026)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
di: Denisov-Blanch, Yegor, et al.
Pubblicazione: (2026)
di: Denisov-Blanch, Yegor, et al.
Pubblicazione: (2026)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
di: Chen, Edward, et al.
Pubblicazione: (2025)
di: Chen, Edward, et al.
Pubblicazione: (2025)
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
di: Stromberg, Nathan, et al.
Pubblicazione: (2024)
di: Stromberg, Nathan, et al.
Pubblicazione: (2024)
Position: Model Collapse Does Not Mean What You Think
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
Optimization and Generalization Guarantees for Weight Normalization
di: Cisneros-Velarde, Pedro, et al.
Pubblicazione: (2024)
di: Cisneros-Velarde, Pedro, et al.
Pubblicazione: (2024)
Structured Prompts Improve Evaluation of Language Models
di: Aali, Asad, et al.
Pubblicazione: (2025)
di: Aali, Asad, et al.
Pubblicazione: (2025)
Prompt Triage: Structured Optimization Enhances Vision-Language Model Performance on Medical Imaging Benchmarks
di: Singhvi, Arnav, et al.
Pubblicazione: (2025)
di: Singhvi, Arnav, et al.
Pubblicazione: (2025)
ReasonEdit: Editing Vision-Language Models using Human Reasoning
di: Qiu, Jiaxing, et al.
Pubblicazione: (2026)
di: Qiu, Jiaxing, et al.
Pubblicazione: (2026)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
di: Gupta, Isha, et al.
Pubblicazione: (2025)
di: Gupta, Isha, et al.
Pubblicazione: (2025)
The Optimization Paradox in Clinical AI Multi-Agent Systems
di: Bedi, Suhana, et al.
Pubblicazione: (2025)
di: Bedi, Suhana, et al.
Pubblicazione: (2025)
Quantifying Variance in Evaluation Benchmarks
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
di: Zhou, Zhanke, et al.
Pubblicazione: (2025)
di: Zhou, Zhanke, et al.
Pubblicazione: (2025)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
di: Chang, Edward Y.
Pubblicazione: (2026)
di: Chang, Edward Y.
Pubblicazione: (2026)
How RLHF Amplifies Sycophancy
di: Shapira, Itai, et al.
Pubblicazione: (2026)
di: Shapira, Itai, et al.
Pubblicazione: (2026)
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
di: Zhao, Zhenyu, et al.
Pubblicazione: (2026)
di: Zhao, Zhenyu, et al.
Pubblicazione: (2026)
Moral Sycophancy in Vision Language Models
di: Rabby, Shadman, et al.
Pubblicazione: (2026)
di: Rabby, Shadman, et al.
Pubblicazione: (2026)
Large language models in medicine: the potentials and pitfalls
di: Omiye, Jesutofunmi A., et al.
Pubblicazione: (2023)
di: Omiye, Jesutofunmi A., et al.
Pubblicazione: (2023)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
di: Patel, Fagun, et al.
Pubblicazione: (2025)
di: Patel, Fagun, et al.
Pubblicazione: (2025)
Decision from Suboptimal Classifiers: Excess Risk Pre- and Post-Calibration
di: Perez-Lebel, Alexandre, et al.
Pubblicazione: (2025)
di: Perez-Lebel, Alexandre, et al.
Pubblicazione: (2025)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models
di: Li, Jiechen, et al.
Pubblicazione: (2026)
di: Li, Jiechen, et al.
Pubblicazione: (2026)
On Fairness of Low-Rank Adaptation of Large Models
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
Documenti analoghi
-
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025) -
Best Practices for Large Language Models in Radiology
di: Bluethgen, Christian, et al.
Pubblicazione: (2024) -
Stop Automating Peer Review Without Rigorous Evaluation
di: Baumann, Joachim, et al.
Pubblicazione: (2026) -
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
di: Iyer, Laya, et al.
Pubblicazione: (2026) -
A Framework for Objective-Driven Dynamical Stochastic Fields
di: Zhang, Yibo Jacky, et al.
Pubblicazione: (2025)