From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vishwarupe, Varad, Shadbolt, Nigel, Jirotka, Marina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
From Rights to Rites: Expectations Management in Smart-Home AI
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
The Collaboration Gap in Human-AI Work
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
A Transparency Paradox? Investigating the Impact of Explanation Specificity and Autonomous Vehicle Perceptual Inaccuracies on Passengers
von: Omeiza, Daniel, et al.
Veröffentlicht: (2024)
von: Omeiza, Daniel, et al.
Veröffentlicht: (2024)
A Rational Analysis of the Effects of Sycophantic AI
von: Batista, Rafael M., et al.
Veröffentlicht: (2026)
von: Batista, Rafael M., et al.
Veröffentlicht: (2026)
Towards AI Agents for Course Instruction in Higher Education: Early Experiences from the Field
von: Simmhan, Yogesh, et al.
Veröffentlicht: (2025)
von: Simmhan, Yogesh, et al.
Veröffentlicht: (2025)
Sycophantic AI makes human interaction feel more effortful and less satisfying over time
von: Ibrahim, Lujain, et al.
Veröffentlicht: (2026)
von: Ibrahim, Lujain, et al.
Veröffentlicht: (2026)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
Why am I Still Seeing This: Measuring the Effectiveness Of Ad Controls and Explanations in AI-Mediated Ad Targeting Systems
von: Castleman, Jane, et al.
Veröffentlicht: (2024)
von: Castleman, Jane, et al.
Veröffentlicht: (2024)
Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
von: Chandra, Kartik, et al.
Veröffentlicht: (2026)
von: Chandra, Kartik, et al.
Veröffentlicht: (2026)
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
Data over dialogue: Why artificial intelligence is unlikely to humanise medicine
von: Hatherley, Joshua
Veröffentlicht: (2025)
von: Hatherley, Joshua
Veröffentlicht: (2025)
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
von: Ding, Shi, et al.
Veröffentlicht: (2025)
von: Ding, Shi, et al.
Veröffentlicht: (2025)
Farsight: Fostering Responsible AI Awareness During AI Application Prototyping
von: Wang, Zijie J., et al.
Veröffentlicht: (2024)
von: Wang, Zijie J., et al.
Veröffentlicht: (2024)
Generative AI in Medicine
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
A Nested Model for AI Design and Validation
von: Dubey, Akshat, et al.
Veröffentlicht: (2024)
von: Dubey, Akshat, et al.
Veröffentlicht: (2024)
Human-in-the-Loop AI for Cheating Ring Detection
von: Shih, Yong-Siang, et al.
Veröffentlicht: (2024)
von: Shih, Yong-Siang, et al.
Veröffentlicht: (2024)
In defence of post-hoc explanations in medical AI
von: Hatherley, Joshua, et al.
Veröffentlicht: (2025)
von: Hatherley, Joshua, et al.
Veröffentlicht: (2025)
Bringing Generative AI to Adaptive Learning in Education
von: Li, Hang, et al.
Veröffentlicht: (2024)
von: Li, Hang, et al.
Veröffentlicht: (2024)
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
von: Kelly, Christopher, et al.
Veröffentlicht: (2026)
von: Kelly, Christopher, et al.
Veröffentlicht: (2026)
Narrowing Action Choices with AI Improves Human Sequential Decisions
von: Straitouri, Eleni, et al.
Veröffentlicht: (2025)
von: Straitouri, Eleni, et al.
Veröffentlicht: (2025)
High hopes for "Deep Medicine"? AI, economics, and the future of care
von: Sparrow, Robert, et al.
Veröffentlicht: (2025)
von: Sparrow, Robert, et al.
Veröffentlicht: (2025)
Ethics and Technical Aspects of Generative AI Models in Digital Content Creation
von: Karagoz, Atahan
Veröffentlicht: (2024)
von: Karagoz, Atahan
Veröffentlicht: (2024)
Deterministic AI Agent Personality Expression through Standard Psychological Diagnostics
von: Kruijssen, J. M. Diederik, et al.
Veröffentlicht: (2025)
von: Kruijssen, J. M. Diederik, et al.
Veröffentlicht: (2025)
Federated learning, ethics, and the double black box problem in medical AI
von: Hatherley, Joshua, et al.
Veröffentlicht: (2025)
von: Hatherley, Joshua, et al.
Veröffentlicht: (2025)
The Digital Transformation in Health: How AI Can Improve the Performance of Health Systems
von: Periáñez, África, et al.
Veröffentlicht: (2024)
von: Periáñez, África, et al.
Veröffentlicht: (2024)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
von: Lehmann, Matthias, et al.
Veröffentlicht: (2024)
von: Lehmann, Matthias, et al.
Veröffentlicht: (2024)
AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
von: Agarwal, Shyam, et al.
Veröffentlicht: (2025)
von: Agarwal, Shyam, et al.
Veröffentlicht: (2025)
ff4ERA: A new Fuzzy Framework for Ethical Risk Assessment in AI
von: Dyoub, Abeer, et al.
Veröffentlicht: (2025)
von: Dyoub, Abeer, et al.
Veröffentlicht: (2025)
AI's Regimes of Representation: A Community-centered Study of Text-to-Image Models in South Asia
von: Qadri, Rida, et al.
Veröffentlicht: (2023)
von: Qadri, Rida, et al.
Veröffentlicht: (2023)
Laboratory-Scale AI: Open-Weight Models are Competitive with ChatGPT Even in Low-Resource Settings
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
von: Wolfe, Robert, et al.
Veröffentlicht: (2024)
A moving target in AI-assisted decision-making: Dataset shift, model updating, and the problem of update opacity
von: Hatherley, Joshua
Veröffentlicht: (2025)
von: Hatherley, Joshua
Veröffentlicht: (2025)
The Malicious Technical Ecosystem: Exposing Limitations in Technical Governance of AI-Generated Non-Consensual Intimate Images of Adults
von: Ding, Michelle L., et al.
Veröffentlicht: (2025)
von: Ding, Michelle L., et al.
Veröffentlicht: (2025)
Culturally-Attuned Moral Machines: Implicit Learning of Human Value Systems by AI through Inverse Reinforcement Learning
von: Oliveira, Nigini, et al.
Veröffentlicht: (2023)
von: Oliveira, Nigini, et al.
Veröffentlicht: (2023)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
ProgressGym: Alignment with a Millennium of Moral Progress
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2024)
Why Trust in AI May Be Inevitable
von: Truong, Nghi, et al.
Veröffentlicht: (2025)
von: Truong, Nghi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026) -
From Rights to Rites: Expectations Management in Smart-Home AI
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026) -
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026) -
The Collaboration Gap in Human-AI Work
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026) -
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)