Towards medical AI misalignment: a preliminary study
Fuente:
arXiv
Saved in:
| Main Authors: | Puccio, Barbara, Castagna, Federico, Tucker, Allan, Veltri, Pierangelo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024)
by: Kanepajs, Artūrs, et al.
Published: (2024)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
by: Rios-Sialer, Ian
Published: (2026)
by: Rios-Sialer, Ian
Published: (2026)
Stars, Stripes, and Silicon: Unravelling the ChatGPT's All-American, Monochrome, Cis-centric Bias
by: Torrielli, Federico
Published: (2024)
by: Torrielli, Federico
Published: (2024)
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying
by: Castagna, Federico, et al.
Published: (2024)
by: Castagna, Federico, et al.
Published: (2024)
Can formal argumentative reasoning enhance LLMs performances?
by: Castagna, Federico, et al.
Published: (2024)
by: Castagna, Federico, et al.
Published: (2024)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
AI Awareness
by: Li, Xiaojian, et al.
Published: (2025)
by: Li, Xiaojian, et al.
Published: (2025)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
by: Xu, Shanshan, et al.
Published: (2025)
by: Xu, Shanshan, et al.
Published: (2025)
Position: The Current AI Conference Model is Unsustainable! Diagnosing the Crisis of Centralized AI Conference
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
AI Literacy in Low-Resource Languages:Insights from creating AI in Yoruba videos
by: Oyewusi, Wuraola
Published: (2024)
by: Oyewusi, Wuraola
Published: (2024)
"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
by: Zhou, Kaitlyn, et al.
Published: (2026)
by: Zhou, Kaitlyn, et al.
Published: (2026)
Human-AI Collaboration or Academic Misconduct? Measuring AI Use in Student Writing Through Stylometric Evidence
by: Oliveira, Eduardo Araujo, et al.
Published: (2025)
by: Oliveira, Eduardo Araujo, et al.
Published: (2025)
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
by: Najjar, Ayat A., et al.
Published: (2025)
by: Najjar, Ayat A., et al.
Published: (2025)
DeepTutor: Towards Agentic Personalized Tutoring
by: Zhao, Bingxi, et al.
Published: (2026)
by: Zhao, Bingxi, et al.
Published: (2026)
Evaluation Framework for AI Systems in "the Wild"
by: Jabbour, Sarah, et al.
Published: (2025)
by: Jabbour, Sarah, et al.
Published: (2025)
Self-Explanation in Social AI Agents
by: Basappa, Rhea, et al.
Published: (2025)
by: Basappa, Rhea, et al.
Published: (2025)
Commercial Persuasion in AI-Mediated Conversations
by: Salvi, Francesco, et al.
Published: (2026)
by: Salvi, Francesco, et al.
Published: (2026)
Conformity and Social Impact on AI Agents
by: Bellina, Alessandro, et al.
Published: (2026)
by: Bellina, Alessandro, et al.
Published: (2026)
When AI Navigates the Fog of War
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
Responsible AI for Test Equity and Quality: The Duolingo English Test as a Case Study
by: Burstein, Jill, et al.
Published: (2024)
by: Burstein, Jill, et al.
Published: (2024)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025)
by: Dwivedi, Ashutosh, et al.
Published: (2025)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
AI Diffusion in Low Resource Language Countries
by: Misra, Amit, et al.
Published: (2025)
by: Misra, Amit, et al.
Published: (2025)
Explainability and Certification of AI-Generated Educational Assessments
by: Yaacoub, Antoun, et al.
Published: (2026)
by: Yaacoub, Antoun, et al.
Published: (2026)
AI Governance and Accountability: An Analysis of Anthropic's Claude
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
by: Cheng, Myra, et al.
Published: (2024)
by: Cheng, Myra, et al.
Published: (2024)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025)
by: Zou, Andy, et al.
Published: (2025)
Noosemia: toward a Cognitive and Phenomenological Account of Intentionality Attribution in Human-Generative AI Interaction
by: De Santis, Enrico, et al.
Published: (2025)
by: De Santis, Enrico, et al.
Published: (2025)
Towards Weakly-Supervised Hate Speech Classification Across Datasets
by: Jin, Yiping, et al.
Published: (2023)
by: Jin, Yiping, et al.
Published: (2023)
Societal AI Research Has Become Less Interdisciplinary
by: Markus, Dror Kris, et al.
Published: (2025)
by: Markus, Dror Kris, et al.
Published: (2025)
Measuring Political Preferences in AI Systems: An Integrative Approach
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Building Trust: Foundations of Security, Safety and Transparency in AI
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
A Unified Framework to Quantify Cultural Intelligence of AI
by: Dev, Sunipa, et al.
Published: (2026)
by: Dev, Sunipa, et al.
Published: (2026)
The Responsible Development of Automated Student Feedback with Generative AI
by: Lindsay, Euan D, et al.
Published: (2023)
by: Lindsay, Euan D, et al.
Published: (2023)
Aalap: AI Assistant for Legal & Paralegal Functions in India
by: Tiwari, Aman, et al.
Published: (2024)
by: Tiwari, Aman, et al.
Published: (2024)
Measuring Human Contribution in AI-Assisted Content Generation
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
Prompt and Prejudice
by: Berlincioni, Lorenzo, et al.
Published: (2024)
by: Berlincioni, Lorenzo, et al.
Published: (2024)
EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content
by: Bi, Shuzhen, et al.
Published: (2026)
by: Bi, Shuzhen, et al.
Published: (2026)
Similar Items
-
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024) -
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
by: Rios-Sialer, Ian
Published: (2026) -
Stars, Stripes, and Silicon: Unravelling the ChatGPT's All-American, Monochrome, Cis-centric Bias
by: Torrielli, Federico
Published: (2024) -
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying
by: Castagna, Federico, et al.
Published: (2024) -
Can formal argumentative reasoning enhance LLMs performances?
by: Castagna, Federico, et al.
Published: (2024)