SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy
Fuente:
arXiv
Guardado en:
| Autores principales: | Bhalla, Joy, Gligorić, Kristina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
por: Cheng, Myra, et al.
Publicado: (2024)
por: Cheng, Myra, et al.
Publicado: (2024)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
por: Christian, Brian, et al.
Publicado: (2026)
por: Christian, Brian, et al.
Publicado: (2026)
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context
por: Arvin, Chuck
Publicado: (2025)
por: Arvin, Chuck
Publicado: (2025)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
por: Natan, Shahar Ben, et al.
Publicado: (2026)
por: Natan, Shahar Ben, et al.
Publicado: (2026)
Sycophancy Claims about Language Models: The Missing Human-in-the-Loop
por: Batzner, Jan, et al.
Publicado: (2025)
por: Batzner, Jan, et al.
Publicado: (2025)
What can large language models do for sustainable food?
por: Thomas, Anna T., et al.
Publicado: (2025)
por: Thomas, Anna T., et al.
Publicado: (2025)
NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
por: Gligoric, Kristina, et al.
Publicado: (2024)
por: Gligoric, Kristina, et al.
Publicado: (2024)
Counterfactual LLM-based Framework for Measuring Rhetorical Style
por: Qiu, Jingyi, et al.
Publicado: (2025)
por: Qiu, Jingyi, et al.
Publicado: (2025)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
por: Sen, Indira, et al.
Publicado: (2023)
por: Sen, Indira, et al.
Publicado: (2023)
Media Manipulations in the Coverage of Events of the Ukrainian Revolution of Dignity: Historical, Linguistic, and Psychological Approaches
por: Khoma, Ivan, et al.
Publicado: (2024)
por: Khoma, Ivan, et al.
Publicado: (2024)
Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles
por: Abdulsalam, Ramatu Oiza, et al.
Publicado: (2025)
por: Abdulsalam, Ramatu Oiza, et al.
Publicado: (2025)
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
por: Batzner, Jan, et al.
Publicado: (2024)
por: Batzner, Jan, et al.
Publicado: (2024)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
por: Padmakumar, Vishakh, et al.
Publicado: (2026)
por: Padmakumar, Vishakh, et al.
Publicado: (2026)
Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks
por: Mehrotra, Navya, et al.
Publicado: (2026)
por: Mehrotra, Navya, et al.
Publicado: (2026)
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
por: Govindarajan, Venkata S, et al.
Publicado: (2023)
por: Govindarajan, Venkata S, et al.
Publicado: (2023)
Attention to Non-Adopters
por: Zhou, Kaitlyn, et al.
Publicado: (2025)
por: Zhou, Kaitlyn, et al.
Publicado: (2025)
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
por: Cunningham, Eoghan, et al.
Publicado: (2025)
por: Cunningham, Eoghan, et al.
Publicado: (2025)
A Practical Method for Generating String Counterfactuals
por: Avitan, Matan, et al.
Publicado: (2024)
por: Avitan, Matan, et al.
Publicado: (2024)
Mapping the Political Discourse in the Brazilian Chamber of Deputies: A Multi-Faceted Computational Approach
por: Soriano, Flávio, et al.
Publicado: (2026)
por: Soriano, Flávio, et al.
Publicado: (2026)
Linguistic Uncertainty and Engagement in Arabic-Language X (formerly Twitter) Discourse
por: Soufan, Mohamed
Publicado: (2026)
por: Soufan, Mohamed
Publicado: (2026)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
por: Fleisig, Eve, et al.
Publicado: (2024)
por: Fleisig, Eve, et al.
Publicado: (2024)
Large Models of What? Mistaking Engineering Achievements for Human Linguistic Agency
por: Birhane, Abeba, et al.
Publicado: (2024)
por: Birhane, Abeba, et al.
Publicado: (2024)
Linguistic Uncertainty and Reply Engagement on X: A Cross-Domain Replication of the Uncertainty-Reply Asymmetry
por: Soufan, Mohamed
Publicado: (2026)
por: Soufan, Mohamed
Publicado: (2026)
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
por: Maurer, Maximilian, et al.
Publicado: (2026)
por: Maurer, Maximilian, et al.
Publicado: (2026)
LinGO: A Linguistic Graph Optimization Framework with LLMs for Interpreting Intents of Online Uncivil Discourse
por: Zhang, Yuan, et al.
Publicado: (2026)
por: Zhang, Yuan, et al.
Publicado: (2026)
Othering and low status framing of immigrant cuisines in US restaurant reviews and large language models
por: Luo, Yiwei, et al.
Publicado: (2023)
por: Luo, Yiwei, et al.
Publicado: (2023)
InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
por: Wan, Yixin, et al.
Publicado: (2025)
por: Wan, Yixin, et al.
Publicado: (2025)
Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data
por: Yamada, Masaru
Publicado: (2026)
por: Yamada, Masaru
Publicado: (2026)
Large Language Models and Forensic Linguistics: Navigating Opportunities and Threats in the Age of Generative AI
por: Mikros, George
Publicado: (2025)
por: Mikros, George
Publicado: (2025)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
por: Liu, Naiming, et al.
Publicado: (2024)
por: Liu, Naiming, et al.
Publicado: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
por: Chien, Jennifer, et al.
Publicado: (2024)
por: Chien, Jennifer, et al.
Publicado: (2024)
LAGAN: Deep Semi-Supervised Linguistic-Anthropology Classification with Conditional Generative Adversarial Neural Network
por: Kamal, Rossi
Publicado: (2023)
por: Kamal, Rossi
Publicado: (2023)
In-Group Love, Out-Group Hate: A Framework to Measure Affective Polarization via Contentious Online Discussions
por: Nettasinghe, Buddhika, et al.
Publicado: (2024)
por: Nettasinghe, Buddhika, et al.
Publicado: (2024)
Mind the Gap: Assessing Wiktionary's Crowd-Sourced Linguistic Knowledge on Morphological Gaps in Two Related Languages
por: Sakunkoo, Jonathan, et al.
Publicado: (2025)
por: Sakunkoo, Jonathan, et al.
Publicado: (2025)
Domain-Independent Deception: A New Taxonomy and Linguistic Analysis
por: Verma, Rakesh M., et al.
Publicado: (2024)
por: Verma, Rakesh M., et al.
Publicado: (2024)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
por: Dorn, Rebecca, et al.
Publicado: (2024)
por: Dorn, Rebecca, et al.
Publicado: (2024)
AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development
por: Šekrst, Kristina, et al.
Publicado: (2024)
por: Šekrst, Kristina, et al.
Publicado: (2024)
Gendered Divides in Online Discussions about Reproductive Rights
por: Rao, Ashwin, et al.
Publicado: (2025)
por: Rao, Ashwin, et al.
Publicado: (2025)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
Ejemplares similares
-
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
por: Cheng, Myra, et al.
Publicado: (2024) -
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
por: Christian, Brian, et al.
Publicado: (2026) -
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context
por: Arvin, Chuck
Publicado: (2025) -
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
por: Natan, Shahar Ben, et al.
Publicado: (2026) -
Sycophancy Claims about Language Models: The Missing Human-in-the-Loop
por: Batzner, Jan, et al.
Publicado: (2025)