Clinical knowledge in LLMs does not translate to human interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Bean, Andrew M., Payne, Rebecca, Parsons, Guy, Kirk, Hannah Rose, Ciro, Juan, Mosquera, Rafael, Monsalve, Sara Hincapié, Ekanayaka, Aruna S., Tarassenko, Lionel, Rocher, Luc, Mahdi, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing and modeling harms from interactions with design patterns in AI interfaces
by: Ibrahim, Lujain, et al.
Published: (2024)
by: Ibrahim, Lujain, et al.
Published: (2024)
Sycophantic AI makes human interaction feel more effortful and less satisfying over time
by: Ibrahim, Lujain, et al.
Published: (2026)
by: Ibrahim, Lujain, et al.
Published: (2026)
SleepVST: Sleep Staging from Near-Infrared Video Signals using Pre-Trained Transformers
by: Carter, Jonathan F., et al.
Published: (2024)
by: Carter, Jonathan F., et al.
Published: (2024)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
by: Khandelwal, Khyati, et al.
Published: (2023)
by: Khandelwal, Khyati, et al.
Published: (2023)
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
by: Kirk, Hannah Rose, et al.
Published: (2024)
by: Kirk, Hannah Rose, et al.
Published: (2024)
Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
by: Zaccour, Juliette, et al.
Published: (2025)
by: Zaccour, Juliette, et al.
Published: (2025)
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
by: Hafner, Franziska Sofia, et al.
Published: (2025)
by: Hafner, Franziska Sofia, et al.
Published: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Training language models to be warm and empathetic makes them less reliable and more sycophantic
by: Ibrahim, Lujain, et al.
Published: (2025)
by: Ibrahim, Lujain, et al.
Published: (2025)
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
Framing Migration: A Computational Analysis of UK Parliamentary Discourse
by: Ghafouri, Vahid, et al.
Published: (2025)
by: Ghafouri, Vahid, et al.
Published: (2025)
Into the crossfire: evaluating the use of a language model to crowdsource gun violence reports
by: Belisario, Adriano, et al.
Published: (2024)
by: Belisario, Adriano, et al.
Published: (2024)
wav2sleep: A Unified Multi-Modal Approach to Sleep Stage Classification from Physiological Signals
by: Carter, Jonathan F., et al.
Published: (2024)
by: Carter, Jonathan F., et al.
Published: (2024)
Conversational AI increases political knowledge as effectively as self-directed internet search
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
by: Röttger, Paul, et al.
Published: (2026)
by: Röttger, Paul, et al.
Published: (2026)
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
by: Bean, Andrew M., et al.
Published: (2024)
by: Bean, Andrew M., et al.
Published: (2024)
Retrieval-augmented reasoning with lean language models
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
People readily follow personal advice from AI but it does not improve their well-being
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Review of multimodal machine learning approaches in healthcare
by: Krones, Felix, et al.
Published: (2024)
by: Krones, Felix, et al.
Published: (2024)
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
by: Padmakumar, Vishakh, et al.
Published: (2024)
by: Padmakumar, Vishakh, et al.
Published: (2024)
Evaluating the role of `Constitutions' for learning from AI feedback
by: Redgate, Saskia, et al.
Published: (2024)
by: Redgate, Saskia, et al.
Published: (2024)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
by: Vidgen, Bertie, et al.
Published: (2023)
by: Vidgen, Bertie, et al.
Published: (2023)
The AI Community Building the Future? A Quantitative Analysis of Development Activity on Hugging Face Hub
by: Osborne, Cailean, et al.
Published: (2024)
by: Osborne, Cailean, et al.
Published: (2024)
Algorithmic Mirror: Designing an Interactive Tool to Promote Self-Reflection for YouTube Recommendations
by: Kondo, Yui, et al.
Published: (2025)
by: Kondo, Yui, et al.
Published: (2025)
Domain adapted machine translation: What does catastrophic forgetting forget and why?
by: Saunders, Danielle, et al.
Published: (2024)
by: Saunders, Danielle, et al.
Published: (2024)
What does it take to get state of the art in simultaneous speech-to-speech translation?
by: Wilmet, Vincent, et al.
Published: (2024)
by: Wilmet, Vincent, et al.
Published: (2024)
The two-way knowledge interaction interface between humans and neural networks
by: He, Zhanliang, et al.
Published: (2024)
by: He, Zhanliang, et al.
Published: (2024)
Evaluating Fine-Tuning Efficiency of Human-Inspired Learning Strategies in Medical Question Answering
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
Contextual effects of sentiment deployment in human and machine translation
by: Comstock, Lindy, et al.
Published: (2025)
by: Comstock, Lindy, et al.
Published: (2025)
Secure Attestation and Dynamic Load Balancing (SALB) for Optimized Container Management: Ensuring Integrity and Enhancing Resource Efficiency
by: K. Aruna
Published: (2025)
by: K. Aruna
Published: (2025)
LLMs in social services: How does chatbot accuracy affect human accuracy?
by: Gosciak, Jennah, et al.
Published: (2026)
by: Gosciak, Jennah, et al.
Published: (2026)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
by: Kirk, Hannah Rose, et al.
Published: (2026)
by: Kirk, Hannah Rose, et al.
Published: (2026)
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
by: Quaye, Jessica, et al.
Published: (2024)
by: Quaye, Jessica, et al.
Published: (2024)
Do Large Language Models have Shared Weaknesses in Medical Question Answering?
by: Bean, Andrew M., et al.
Published: (2023)
by: Bean, Andrew M., et al.
Published: (2023)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
An analysis of AI Decision under Risk: Prospect theory emerges in Large Language Models
by: Payne, Kenneth
Published: (2025)
by: Payne, Kenneth
Published: (2025)
Towards interactive evaluations for interaction harms in human-AI systems
by: Ibrahim, Lujain, et al.
Published: (2024)
by: Ibrahim, Lujain, et al.
Published: (2024)
Can Large Language Models Simulate Human Responses? A Case Study of Stated Preference Experiments in the Context of Heating-related Choices
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Similar Items
-
Characterizing and modeling harms from interactions with design patterns in AI interfaces
by: Ibrahim, Lujain, et al.
Published: (2024) -
Sycophantic AI makes human interaction feel more effortful and less satisfying over time
by: Ibrahim, Lujain, et al.
Published: (2026) -
SleepVST: Sleep Staging from Near-Infrared Video Signals using Pre-Trained Transformers
by: Carter, Jonathan F., et al.
Published: (2024) -
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
by: Khandelwal, Khyati, et al.
Published: (2023) -
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
by: Kirk, Hannah Rose, et al.
Published: (2024)