Confirmation bias: A challenge for scalable oversight
Fuente:
arXiv
Saved in:
| Main Authors: | Recchia, Gabriel, Mangat, Chatrik Singh, Nyachhyon, Jinu, Sharma, Mridul, Canavan, Callum, Epstein-Gross, Dylan, Abdulbari, Muhammed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025)
by: Recchia, Gabriel, et al.
Published: (2025)
Towards physician-centered oversight of conversational diagnostic AI
by: Vedadi, Elahe, et al.
Published: (2025)
by: Vedadi, Elahe, et al.
Published: (2025)
Confirmation Bias in Generative AI Chatbots: Mechanisms, Risks, Mitigation Strategies, and Future Research Directions
by: Du, Yiran
Published: (2025)
by: Du, Yiran
Published: (2025)
AI Standardized Patient Improves Human Conversations in Advanced Cancer Care
by: Haut, Kurtis, et al.
Published: (2025)
by: Haut, Kurtis, et al.
Published: (2025)
Creating Disability Story Videos with Generative AI: Motivation, Expression, and Sharing
by: Niu, Shuo, et al.
Published: (2026)
by: Niu, Shuo, et al.
Published: (2026)
To Bias or Not to Bias: Detecting bias in News with bias-detector
by: Ghosh, Himel, et al.
Published: (2025)
by: Ghosh, Himel, et al.
Published: (2025)
"Even explanations will not help in trusting [this] fundamentally biased system": A Predictive Policing Case-Study
by: Mehrotra, Siddharth, et al.
Published: (2025)
by: Mehrotra, Siddharth, et al.
Published: (2025)
Community-Centered Spatial Intelligence for Climate Adaptation at Nova Scotia's Eastern Shore
by: Spadon, Gabriel, et al.
Published: (2025)
by: Spadon, Gabriel, et al.
Published: (2025)
On using AI for EEG-based BCI applications: problems, current challenges and future trends
by: Barbera, Thomas, et al.
Published: (2025)
by: Barbera, Thomas, et al.
Published: (2025)
Predicting the usability of mobile applications using AI tools: the rise of large user interface models, opportunities, and challenges
by: Namoun, Abdallah, et al.
Published: (2024)
by: Namoun, Abdallah, et al.
Published: (2024)
AnnoSense: A Framework for Physiological Emotion Data Collection in Everyday Settings for AI
by: Singh, Pragya, et al.
Published: (2025)
by: Singh, Pragya, et al.
Published: (2025)
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
by: Yu, Xiaofan, et al.
Published: (2025)
by: Yu, Xiaofan, et al.
Published: (2025)
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024)
by: Shaikh, Ammar, et al.
Published: (2024)
Can AI agents understand spoken conversations about data visualizations in online meetings?
by: Sharma, Rizul, et al.
Published: (2025)
by: Sharma, Rizul, et al.
Published: (2025)
Toward Safe Evolution of Artificial Intelligence (AI) based Conversational Agents to Support Adolescent Mental and Sexual Health Knowledge Discovery
by: Park, Jinkyung, et al.
Published: (2024)
by: Park, Jinkyung, et al.
Published: (2024)
Leveraging Large Language Models (LLMs) to Support Collaborative Human-AI Online Risk Data Annotation
by: Park, Jinkyung, et al.
Published: (2024)
by: Park, Jinkyung, et al.
Published: (2024)
AI Autonomy Coefficient ($α$): Defining Boundaries for Responsible AI Systems
by: Mairittha, Nattaya, et al.
Published: (2025)
by: Mairittha, Nattaya, et al.
Published: (2025)
Boli: A dataset for understanding stuttering experience and analyzing stuttered speech
by: Batra, Ashita, et al.
Published: (2025)
by: Batra, Ashita, et al.
Published: (2025)
Impact of Multimodal and Conversational AI on Learning Outcomes and Experience
by: Taneja, Karan, et al.
Published: (2026)
by: Taneja, Karan, et al.
Published: (2026)
Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
by: Chappidi, Shreya, et al.
Published: (2026)
by: Chappidi, Shreya, et al.
Published: (2026)
PersonaFlow: Designing LLM-Simulated Expert Perspectives for Enhanced Research Ideation
by: Liu, Yiren, et al.
Published: (2024)
by: Liu, Yiren, et al.
Published: (2024)
Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
by: Wang, Qiaosi, et al.
Published: (2026)
by: Wang, Qiaosi, et al.
Published: (2026)
Toward a Human-Centered AI-assisted Colonoscopy System in Australia
by: Chen, Hsiang-Ting, et al.
Published: (2025)
by: Chen, Hsiang-Ting, et al.
Published: (2025)
AI-powered virtual eye: perspective, challenges and opportunities
by: Wu, Yue, et al.
Published: (2025)
by: Wu, Yue, et al.
Published: (2025)
Toward Human-AI Alignment in Large-Scale Multi-Player Games
by: Sharma, Sugandha, et al.
Published: (2024)
by: Sharma, Sugandha, et al.
Published: (2024)
From Evidence to Decision: Exploring Evaluative AI
by: Le, Thao, et al.
Published: (2024)
by: Le, Thao, et al.
Published: (2024)
Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study
by: Li, Harry, et al.
Published: (2025)
by: Li, Harry, et al.
Published: (2025)
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
Experiential Explanations for Reinforcement Learning
by: Alabdulkarim, Amal, et al.
Published: (2022)
by: Alabdulkarim, Amal, et al.
Published: (2022)
Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video
by: Groh, Matthew, et al.
Published: (2022)
by: Groh, Matthew, et al.
Published: (2022)
Crowdsourced human-based computational approach for tagging peripheral blood smear sample images from Sickle Cell Disease patients using non-expert users
by: Rubio, José María Buades, et al.
Published: (2025)
by: Rubio, José María Buades, et al.
Published: (2025)
Understanding Critical Thinking in Generative Artificial Intelligence Use: Development, Validation, and Correlates of the Critical Thinking in AI Use Scale
by: Lau, Gabriel R., et al.
Published: (2025)
by: Lau, Gabriel R., et al.
Published: (2025)
Evaluating AI Alignment in LLMs: Output Analysis of Value Priorities Across 75 Models with Human Benchmarking
by: Lau, Gabriel Rongyang, et al.
Published: (2025)
by: Lau, Gabriel Rongyang, et al.
Published: (2025)
Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals
by: Ide, Ayae, et al.
Published: (2025)
by: Ide, Ayae, et al.
Published: (2025)
TactStyle: Generating Tactile Textures with Generative AI for Digital Fabrication
by: Faruqi, Faraz, et al.
Published: (2025)
by: Faruqi, Faraz, et al.
Published: (2025)
Exploring the Requirements of Clinicians for Explainable AI Decision Support Systems in Intensive Care
by: Clark, Jeffrey N., et al.
Published: (2024)
by: Clark, Jeffrey N., et al.
Published: (2024)
Data Ethics Emergency Drill: A Toolbox for Discussing Responsible AI for Industry Teams
by: Hanschke, Vanessa Aisyahsari, et al.
Published: (2024)
by: Hanschke, Vanessa Aisyahsari, et al.
Published: (2024)
Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions
by: Sharma, Paras, et al.
Published: (2026)
by: Sharma, Paras, et al.
Published: (2026)
XAIxArts Manifesto: Explainable AI for the Arts
by: Bryan-Kinns, Nick, et al.
Published: (2025)
by: Bryan-Kinns, Nick, et al.
Published: (2025)
The Odyssey of the Fittest: Can Agents Survive and Still Be Good?
by: Waldner, Dylan, et al.
Published: (2025)
by: Waldner, Dylan, et al.
Published: (2025)
Similar Items
-
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025) -
Towards physician-centered oversight of conversational diagnostic AI
by: Vedadi, Elahe, et al.
Published: (2025) -
Confirmation Bias in Generative AI Chatbots: Mechanisms, Risks, Mitigation Strategies, and Future Research Directions
by: Du, Yiran
Published: (2025) -
AI Standardized Patient Improves Human Conversations in Advanced Cancer Care
by: Haut, Kurtis, et al.
Published: (2025) -
Creating Disability Story Videos with Generative AI: Motivation, Expression, and Sharing
by: Niu, Shuo, et al.
Published: (2026)