BASIL: Bayesian Assessment of Sycophancy in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Atwell, Katherine, Heydari, Pedram, Sicilia, Anthony, Alikhani, Malihe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026)
von: Imai, Saki, et al.
Veröffentlicht: (2026)
Accounting for Sycophancy in Language Model Uncertainty Estimation
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Generating Signed Language Instructions in Large-Scale Dialogue Systems
von: İnan, Mert, et al.
Veröffentlicht: (2024)
von: İnan, Mert, et al.
Veröffentlicht: (2024)
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI
von: Asano, Yuya, et al.
Veröffentlicht: (2025)
von: Asano, Yuya, et al.
Veröffentlicht: (2025)
HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations
von: Sicilia, Anthony, et al.
Veröffentlicht: (2023)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2023)
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
von: Imai, Saki, et al.
Veröffentlicht: (2025)
von: Imai, Saki, et al.
Veröffentlicht: (2025)
Measuring How (Not Just Whether) VLMs Build Common Ground
von: Imai, Saki, et al.
Veröffentlicht: (2025)
von: Imai, Saki, et al.
Veröffentlicht: (2025)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Studying and Mitigating Biases in Sign Language Understanding Models
von: Atwell, Katherine, et al.
Veröffentlicht: (2024)
von: Atwell, Katherine, et al.
Veröffentlicht: (2024)
An Active Learning Framework for Inclusive Generation by Large Language Models
von: Hassan, Sabit, et al.
Veröffentlicht: (2024)
von: Hassan, Sabit, et al.
Veröffentlicht: (2024)
Active Learning for Robust and Representative LLM Generation in Safety-Critical Scenarios
von: Hassan, Sabit, et al.
Veröffentlicht: (2024)
von: Hassan, Sabit, et al.
Veröffentlicht: (2024)
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
von: İnan, Mert, et al.
Veröffentlicht: (2025)
von: İnan, Mert, et al.
Veröffentlicht: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
von: Zhang, Chaowei, et al.
Veröffentlicht: (2026)
von: Zhang, Chaowei, et al.
Veröffentlicht: (2026)
Modeling Intensification for Sign Language Generation: A Computational Approach
von: İnan, Mert, et al.
Veröffentlicht: (2022)
von: İnan, Mert, et al.
Veröffentlicht: (2022)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
Sycophancy Hides Linearly in the Attention Heads
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
How people talk about each other: Modeling Generalized Intergroup Bias and Emotion
von: Govindarajan, Venkata S, et al.
Veröffentlicht: (2022)
von: Govindarajan, Venkata S, et al.
Veröffentlicht: (2022)
Sycophancy in Large Language Models: Causes and Mitigations
von: Malmqvist, Lars
Veröffentlicht: (2024)
von: Malmqvist, Lars
Veröffentlicht: (2024)
Ta'keed: The First Generative Fact-Checking System for Arabic Claims
von: Althabiti, Saud, et al.
Veröffentlicht: (2024)
von: Althabiti, Saud, et al.
Veröffentlicht: (2024)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
von: Guo, Kevin H., et al.
Veröffentlicht: (2026)
von: Guo, Kevin H., et al.
Veröffentlicht: (2026)
"Nothing about us without us": Perspectives of Global Deaf and Hard-of-hearing Community Members on Sign Language Technologies
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
von: Denison, Carson, et al.
Veröffentlicht: (2024)
von: Denison, Carson, et al.
Veröffentlicht: (2024)
Change My View? The Dynamics of Persuasion and Polarization in Online Discourse
von: Freeborn, David, et al.
Veröffentlicht: (2026)
von: Freeborn, David, et al.
Veröffentlicht: (2026)
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
von: Jain, Raunak
Veröffentlicht: (2026)
von: Jain, Raunak
Veröffentlicht: (2026)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
von: Pandey, Sanskar, et al.
Veröffentlicht: (2025)
von: Pandey, Sanskar, et al.
Veröffentlicht: (2025)
MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
von: Hu, Jingyu, et al.
Veröffentlicht: (2025)
von: Hu, Jingyu, et al.
Veröffentlicht: (2025)
Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework
von: Zhao, Yunpu, et al.
Veröffentlicht: (2024)
von: Zhao, Yunpu, et al.
Veröffentlicht: (2024)
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
von: Shah, Arya, et al.
Veröffentlicht: (2026)
von: Shah, Arya, et al.
Veröffentlicht: (2026)
Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks
von: Han, Tianze, et al.
Veröffentlicht: (2026)
von: Han, Tianze, et al.
Veröffentlicht: (2026)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
von: Christian, Brian, et al.
Veröffentlicht: (2026)
von: Christian, Brian, et al.
Veröffentlicht: (2026)
PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
von: Çelebi, Yusuf, et al.
Veröffentlicht: (2025)
von: Çelebi, Yusuf, et al.
Veröffentlicht: (2025)
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
von: Kasprova, Vira, et al.
Veröffentlicht: (2026)
von: Kasprova, Vira, et al.
Veröffentlicht: (2026)
Towards Understanding Sycophancy in Language Models
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
FarsiMCQGen: a Persian Multiple-choice Question Generation Framework
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2025)
von: Rad, Mohammad Heydari, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026) -
Accounting for Sycophancy in Language Model Uncertainty Estimation
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024) -
Generating Signed Language Instructions in Large-Scale Dialogue Systems
von: İnan, Mert, et al.
Veröffentlicht: (2024) -
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024) -
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)