Gespeichert in:
| Hauptverfasser: | Shapira, Itai, Benade, Gerdus, Procaccia, Ariel D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.01002 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pairwise Calibrated Rewards for Pluralistic Alignment
von: Halpern, Daniel, et al.
Veröffentlicht: (2025)
von: Halpern, Daniel, et al.
Veröffentlicht: (2025)
Axioms for AI Alignment from Human Feedback
von: Ge, Luise, et al.
Veröffentlicht: (2024)
von: Ge, Luise, et al.
Veröffentlicht: (2024)
Generative Social Choice
von: Fish, Sara, et al.
Veröffentlicht: (2023)
von: Fish, Sara, et al.
Veröffentlicht: (2023)
Offline Local Search for Online Stochastic Bandits
von: Benadè, Gerdus, et al.
Veröffentlicht: (2026)
von: Benadè, Gerdus, et al.
Veröffentlicht: (2026)
Incentives in Federated Learning with Heterogeneous Agents
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
Embeddings for Preferences, Not Semantics
von: Blair, Carter, et al.
Veröffentlicht: (2026)
von: Blair, Carter, et al.
Veröffentlicht: (2026)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
Learning Social Welfare Functions
von: Pardeshi, Kanad Shrikar, et al.
Veröffentlicht: (2024)
von: Pardeshi, Kanad Shrikar, et al.
Veröffentlicht: (2024)
Generative Social Choice: The Next Generation
von: Boehmer, Niclas, et al.
Veröffentlicht: (2025)
von: Boehmer, Niclas, et al.
Veröffentlicht: (2025)
Clone-Robust AI Alignment
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
Bias Detection Via Signaling
von: Chen, Yiling, et al.
Veröffentlicht: (2024)
von: Chen, Yiling, et al.
Veröffentlicht: (2024)
Policy Aggregation
von: Alamdari, Parand A., et al.
Veröffentlicht: (2024)
von: Alamdari, Parand A., et al.
Veröffentlicht: (2024)
Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies
von: Assos, Angelos, et al.
Veröffentlicht: (2025)
von: Assos, Angelos, et al.
Veröffentlicht: (2025)
Strategic Classification With Externalities
von: Hossain, Safwan, et al.
Veröffentlicht: (2024)
von: Hossain, Safwan, et al.
Veröffentlicht: (2024)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Finding Common Ground in a Sea of Alternatives
von: Chooi, Jay, et al.
Veröffentlicht: (2026)
von: Chooi, Jay, et al.
Veröffentlicht: (2026)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
Moral Sycophancy in Vision Language Models
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
SycEval: Evaluating LLM Sycophancy
von: Fanous, Aaron, et al.
Veröffentlicht: (2025)
von: Fanous, Aaron, et al.
Veröffentlicht: (2025)
Question the Questions: Auditing Representation in Online Deliberative Processes
von: De, Soham, et al.
Veröffentlicht: (2025)
von: De, Soham, et al.
Veröffentlicht: (2025)
Linear Probe Penalties Reduce LLM Sycophancy
von: Papadatos, Henry, et al.
Veröffentlicht: (2024)
von: Papadatos, Henry, et al.
Veröffentlicht: (2024)
When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models
von: Li, Jiechen, et al.
Veröffentlicht: (2026)
von: Li, Jiechen, et al.
Veröffentlicht: (2026)
Direct Alignment with Heterogeneous Preferences
von: Shirali, Ali, et al.
Veröffentlicht: (2025)
von: Shirali, Ali, et al.
Veröffentlicht: (2025)
Sycophancy Hides Linearly in the Attention Heads
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
Adaptive Contracts for Cost-Effective AI Delegation
von: Saig, Eden, et al.
Veröffentlicht: (2026)
von: Saig, Eden, et al.
Veröffentlicht: (2026)
Sycophancy in Large Language Models: Causes and Mitigations
von: Malmqvist, Lars
Veröffentlicht: (2024)
von: Malmqvist, Lars
Veröffentlicht: (2024)
Consistency Training Helps Stop Sycophancy and Jailbreaks
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
von: Maltbie, Benjamin, et al.
Veröffentlicht: (2026)
von: Maltbie, Benjamin, et al.
Veröffentlicht: (2026)
Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
von: Törnberg, Petter, et al.
Veröffentlicht: (2026)
von: Törnberg, Petter, et al.
Veröffentlicht: (2026)
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2025)
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
von: Chang, Edward Y.
Veröffentlicht: (2026)
von: Chang, Edward Y.
Veröffentlicht: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
ROCM: RLHF on consistency models
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2025)
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pairwise Calibrated Rewards for Pluralistic Alignment
von: Halpern, Daniel, et al.
Veröffentlicht: (2025) -
Axioms for AI Alignment from Human Feedback
von: Ge, Luise, et al.
Veröffentlicht: (2024) -
Generative Social Choice
von: Fish, Sara, et al.
Veröffentlicht: (2023) -
Offline Local Search for Online Stochastic Bandits
von: Benadè, Gerdus, et al.
Veröffentlicht: (2026) -
Incentives in Federated Learning with Heterogeneous Agents
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)