Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
Fuente:
arXiv
Salvato in:
| Autori principali: | Törnberg, Petter, Schimmel, Michelle |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
di: Törnberg, Petter
Pubblicazione: (2026)
di: Törnberg, Petter
Pubblicazione: (2026)
Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation
di: Pagan, Nicolò, et al.
Pubblicazione: (2026)
di: Pagan, Nicolò, et al.
Pubblicazione: (2026)
Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
di: Larooij, Maik, et al.
Pubblicazione: (2025)
di: Larooij, Maik, et al.
Pubblicazione: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
Capturing Bias Diversity in LLMs
di: Gosavi, Purva Prasad, et al.
Pubblicazione: (2024)
di: Gosavi, Purva Prasad, et al.
Pubblicazione: (2024)
BASIL: Bayesian Assessment of Sycophancy in LLMs
di: Atwell, Katherine, et al.
Pubblicazione: (2025)
di: Atwell, Katherine, et al.
Pubblicazione: (2025)
A Scalable Entity-Based Framework for Auditing Bias in LLMs
di: Elbouanani, Akram, et al.
Pubblicazione: (2026)
di: Elbouanani, Akram, et al.
Pubblicazione: (2026)
Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks
di: Han, Tianze, et al.
Pubblicazione: (2026)
di: Han, Tianze, et al.
Pubblicazione: (2026)
How RLHF Amplifies Sycophancy
di: Shapira, Itai, et al.
Pubblicazione: (2026)
di: Shapira, Itai, et al.
Pubblicazione: (2026)
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
di: Zhang, Chaowei, et al.
Pubblicazione: (2026)
di: Zhang, Chaowei, et al.
Pubblicazione: (2026)
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
di: Kasprova, Vira, et al.
Pubblicazione: (2026)
di: Kasprova, Vira, et al.
Pubblicazione: (2026)
Elite Political Discourse has Become More Toxic in Western Countries
di: Törnberg, Petter, et al.
Pubblicazione: (2025)
di: Törnberg, Petter, et al.
Pubblicazione: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
Moral Sycophancy in Vision Language Models
di: Rabby, Shadman, et al.
Pubblicazione: (2026)
di: Rabby, Shadman, et al.
Pubblicazione: (2026)
SycEval: Evaluating LLM Sycophancy
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
di: Zhou, Wenrui, et al.
Pubblicazione: (2025)
di: Zhou, Wenrui, et al.
Pubblicazione: (2025)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
di: Kasneci, Enkelejda, et al.
Pubblicazione: (2026)
di: Kasneci, Enkelejda, et al.
Pubblicazione: (2026)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
di: Natan, Shahar Ben, et al.
Pubblicazione: (2026)
di: Natan, Shahar Ben, et al.
Pubblicazione: (2026)
Linear Probe Penalties Reduce LLM Sycophancy
di: Papadatos, Henry, et al.
Pubblicazione: (2024)
di: Papadatos, Henry, et al.
Pubblicazione: (2024)
When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models
di: Li, Jiechen, et al.
Pubblicazione: (2026)
di: Li, Jiechen, et al.
Pubblicazione: (2026)
The City as an Anti‐Growth Machine
di: Petter Törnberg
Pubblicazione: (2026)
di: Petter Törnberg
Pubblicazione: (2026)
Shifts in U.S. Social Media Use, 2020-2024: Decline, Fragmentation, and Enduring Polarization
di: Törnberg, Petter
Pubblicazione: (2025)
di: Törnberg, Petter
Pubblicazione: (2025)
Best Practices for Text Annotation with Large Language Models
di: Törnberg, Petter
Pubblicazione: (2024)
di: Törnberg, Petter
Pubblicazione: (2024)
Online Homogeneity Can Emerge Without Filtering Algorithms or Homophily Preferences
di: Törnberg, Petter
Pubblicazione: (2025)
di: Törnberg, Petter
Pubblicazione: (2025)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
di: DiGiuseppe, Matthew, et al.
Pubblicazione: (2026)
di: DiGiuseppe, Matthew, et al.
Pubblicazione: (2026)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
di: Sakhawat, Adib, et al.
Pubblicazione: (2026)
di: Sakhawat, Adib, et al.
Pubblicazione: (2026)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2026)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2026)
Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification
di: Elbouanani, Akram, et al.
Pubblicazione: (2025)
di: Elbouanani, Akram, et al.
Pubblicazione: (2025)
Political Bias in LLMs: Unaligned Moral Values in Agent-centric Simulations
di: Münker, Simon
Pubblicazione: (2024)
di: Münker, Simon
Pubblicazione: (2024)
Sycophancy Hides Linearly in the Attention Heads
di: Genadi, Rifo, et al.
Pubblicazione: (2026)
di: Genadi, Rifo, et al.
Pubblicazione: (2026)
Global AI Bias Audit for Technical Governance
di: Hung, Jason
Pubblicazione: (2026)
di: Hung, Jason
Pubblicazione: (2026)
Who Attacks, and Why? Using LLMs to Identify Negative Campaigning in 18M Tweets across 19 Countries
di: Hartman, Victor, et al.
Pubblicazione: (2025)
di: Hartman, Victor, et al.
Pubblicazione: (2025)
Law and the Emerging Political Economy of Algorithmic Audits
di: Terzis, Petros, et al.
Pubblicazione: (2024)
di: Terzis, Petros, et al.
Pubblicazione: (2024)
Sycophancy in Large Language Models: Causes and Mitigations
di: Malmqvist, Lars
Pubblicazione: (2024)
di: Malmqvist, Lars
Pubblicazione: (2024)
Consistency Training Helps Stop Sycophancy and Jailbreaks
di: Irpan, Alex, et al.
Pubblicazione: (2025)
di: Irpan, Alex, et al.
Pubblicazione: (2025)
When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications
di: Cao, Hongliu, et al.
Pubblicazione: (2026)
di: Cao, Hongliu, et al.
Pubblicazione: (2026)
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
di: Wang, Libo
Pubblicazione: (2024)
di: Wang, Libo
Pubblicazione: (2024)
Documenti analoghi
-
Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
di: Törnberg, Petter
Pubblicazione: (2026) -
Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation
di: Pagan, Nicolò, et al.
Pubblicazione: (2026) -
Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
di: Larooij, Maik, et al.
Pubblicazione: (2025) -
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
di: Barkett, Emilio, et al.
Pubblicazione: (2025) -
Capturing Bias Diversity in LLMs
di: Gosavi, Purva Prasad, et al.
Pubblicazione: (2024)