Guardado en:
| Autores principales: | Shapira, Itai, Benade, Gerdus, Procaccia, Ariel D. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.01002 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pairwise Calibrated Rewards for Pluralistic Alignment
por: Halpern, Daniel, et al.
Publicado: (2025)
por: Halpern, Daniel, et al.
Publicado: (2025)
Axioms for AI Alignment from Human Feedback
por: Ge, Luise, et al.
Publicado: (2024)
por: Ge, Luise, et al.
Publicado: (2024)
Generative Social Choice
por: Fish, Sara, et al.
Publicado: (2023)
por: Fish, Sara, et al.
Publicado: (2023)
Offline Local Search for Online Stochastic Bandits
por: Benadè, Gerdus, et al.
Publicado: (2026)
por: Benadè, Gerdus, et al.
Publicado: (2026)
Incentives in Federated Learning with Heterogeneous Agents
por: Procaccia, Ariel D., et al.
Publicado: (2025)
por: Procaccia, Ariel D., et al.
Publicado: (2025)
Embeddings for Preferences, Not Semantics
por: Blair, Carter, et al.
Publicado: (2026)
por: Blair, Carter, et al.
Publicado: (2026)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
Learning Social Welfare Functions
por: Pardeshi, Kanad Shrikar, et al.
Publicado: (2024)
por: Pardeshi, Kanad Shrikar, et al.
Publicado: (2024)
Generative Social Choice: The Next Generation
por: Boehmer, Niclas, et al.
Publicado: (2025)
por: Boehmer, Niclas, et al.
Publicado: (2025)
Clone-Robust AI Alignment
por: Procaccia, Ariel D., et al.
Publicado: (2025)
por: Procaccia, Ariel D., et al.
Publicado: (2025)
Bias Detection Via Signaling
por: Chen, Yiling, et al.
Publicado: (2024)
por: Chen, Yiling, et al.
Publicado: (2024)
Policy Aggregation
por: Alamdari, Parand A., et al.
Publicado: (2024)
por: Alamdari, Parand A., et al.
Publicado: (2024)
Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies
por: Assos, Angelos, et al.
Publicado: (2025)
por: Assos, Angelos, et al.
Publicado: (2025)
Strategic Classification With Externalities
por: Hossain, Safwan, et al.
Publicado: (2024)
por: Hossain, Safwan, et al.
Publicado: (2024)
How to Evaluate Reward Models for RLHF
por: Frick, Evan, et al.
Publicado: (2024)
por: Frick, Evan, et al.
Publicado: (2024)
Finding Common Ground in a Sea of Alternatives
por: Chooi, Jay, et al.
Publicado: (2026)
por: Chooi, Jay, et al.
Publicado: (2026)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
por: Natan, Shahar Ben, et al.
Publicado: (2026)
por: Natan, Shahar Ben, et al.
Publicado: (2026)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
por: Kasneci, Enkelejda, et al.
Publicado: (2026)
por: Kasneci, Enkelejda, et al.
Publicado: (2026)
Moral Sycophancy in Vision Language Models
por: Rabby, Shadman, et al.
Publicado: (2026)
por: Rabby, Shadman, et al.
Publicado: (2026)
SycEval: Evaluating LLM Sycophancy
por: Fanous, Aaron, et al.
Publicado: (2025)
por: Fanous, Aaron, et al.
Publicado: (2025)
Question the Questions: Auditing Representation in Online Deliberative Processes
por: De, Soham, et al.
Publicado: (2025)
por: De, Soham, et al.
Publicado: (2025)
Linear Probe Penalties Reduce LLM Sycophancy
por: Papadatos, Henry, et al.
Publicado: (2024)
por: Papadatos, Henry, et al.
Publicado: (2024)
When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models
por: Li, Jiechen, et al.
Publicado: (2026)
por: Li, Jiechen, et al.
Publicado: (2026)
Direct Alignment with Heterogeneous Preferences
por: Shirali, Ali, et al.
Publicado: (2025)
por: Shirali, Ali, et al.
Publicado: (2025)
Sycophancy Hides Linearly in the Attention Heads
por: Genadi, Rifo, et al.
Publicado: (2026)
por: Genadi, Rifo, et al.
Publicado: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
por: Atwell, Katherine, et al.
Publicado: (2025)
por: Atwell, Katherine, et al.
Publicado: (2025)
Adaptive Contracts for Cost-Effective AI Delegation
por: Saig, Eden, et al.
Publicado: (2026)
por: Saig, Eden, et al.
Publicado: (2026)
Sycophancy in Large Language Models: Causes and Mitigations
por: Malmqvist, Lars
Publicado: (2024)
por: Malmqvist, Lars
Publicado: (2024)
Consistency Training Helps Stop Sycophancy and Jailbreaks
por: Irpan, Alex, et al.
Publicado: (2025)
por: Irpan, Alex, et al.
Publicado: (2025)
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
por: Maltbie, Benjamin, et al.
Publicado: (2026)
por: Maltbie, Benjamin, et al.
Publicado: (2026)
Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
por: Törnberg, Petter, et al.
Publicado: (2026)
por: Törnberg, Petter, et al.
Publicado: (2026)
SOAP: Improving and Stabilizing Shampoo using Adam
por: Vyas, Nikhil, et al.
Publicado: (2024)
por: Vyas, Nikhil, et al.
Publicado: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024)
por: Dong, Hanze, et al.
Publicado: (2024)
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
por: Wang, Libo
Publicado: (2024)
por: Wang, Libo
Publicado: (2024)
PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models
por: Rahman, A. B. M. Ashikur, et al.
Publicado: (2025)
por: Rahman, A. B. M. Ashikur, et al.
Publicado: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
por: Sahoo, Subramanyam
Publicado: (2026)
por: Sahoo, Subramanyam
Publicado: (2026)
Dataset Reset Policy Optimization for RLHF
por: Chang, Jonathan D., et al.
Publicado: (2024)
por: Chang, Jonathan D., et al.
Publicado: (2024)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
por: Chang, Edward Y.
Publicado: (2026)
por: Chang, Edward Y.
Publicado: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
por: Hu, Jian, et al.
Publicado: (2024)
por: Hu, Jian, et al.
Publicado: (2024)
ROCM: RLHF on consistency models
por: Shekhar, Shivanshu, et al.
Publicado: (2025)
por: Shekhar, Shivanshu, et al.
Publicado: (2025)
Ejemplares similares
-
Pairwise Calibrated Rewards for Pluralistic Alignment
por: Halpern, Daniel, et al.
Publicado: (2025) -
Axioms for AI Alignment from Human Feedback
por: Ge, Luise, et al.
Publicado: (2024) -
Generative Social Choice
por: Fish, Sara, et al.
Publicado: (2023) -
Offline Local Search for Online Stochastic Bandits
por: Benadè, Gerdus, et al.
Publicado: (2026) -
Incentives in Federated Learning with Heterogeneous Agents
por: Procaccia, Ariel D., et al.
Publicado: (2025)