Medical large language models are easily distracted
Fuente:
arXiv
Guardado en:
| Autores principales: | Vishwanath, Krithik, Alyakin, Anton, Alber, Daniel Alexander, Lee, Jin Vivian, Kondziolka, Douglas, Oermann, Eric Karl |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
por: Vishwanath, Krithik, et al.
Publicado: (2025)
por: Vishwanath, Krithik, et al.
Publicado: (2025)
MedMobile: A mobile-sized language model with clinical capabilities
por: Vishwanath, Krithik, et al.
Publicado: (2024)
por: Vishwanath, Krithik, et al.
Publicado: (2024)
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
por: Vishwanath, Krithik, et al.
Publicado: (2025)
por: Vishwanath, Krithik, et al.
Publicado: (2025)
CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications
por: Alyakin, Anton, et al.
Publicado: (2025)
por: Alyakin, Anton, et al.
Publicado: (2025)
Medication counseling with large language models: balancing flexibility and rigidity
por: Sabel, Joar, et al.
Publicado: (2025)
por: Sabel, Joar, et al.
Publicado: (2025)
Automatic deductive coding in discourse analysis: an application of large language models in learning analytics
por: Zhang, Lishan, et al.
Publicado: (2024)
por: Zhang, Lishan, et al.
Publicado: (2024)
Large-scale moral machine experiment on large language models
por: Ahmad, Muhammad Shahrul Zaim bin, et al.
Publicado: (2024)
por: Ahmad, Muhammad Shahrul Zaim bin, et al.
Publicado: (2024)
Addressing cognitive bias in medical language models
por: Schmidgall, Samuel, et al.
Publicado: (2024)
por: Schmidgall, Samuel, et al.
Publicado: (2024)
Creativity Benchmark: A benchmark for marketing creativity for large language models
por: Bhat, Ninad, et al.
Publicado: (2025)
por: Bhat, Ninad, et al.
Publicado: (2025)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
por: Nusrat, Humza, et al.
Publicado: (2025)
por: Nusrat, Humza, et al.
Publicado: (2025)
Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners
por: Vaccaro Jr, Michael, et al.
Publicado: (2024)
por: Vaccaro Jr, Michael, et al.
Publicado: (2024)
The role of large language models in UI/UX design: A systematic literature review
por: Ahmed, Ammar, et al.
Publicado: (2025)
por: Ahmed, Ammar, et al.
Publicado: (2025)
Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models
por: Peters, Uwe, et al.
Publicado: (2026)
por: Peters, Uwe, et al.
Publicado: (2026)
Assessing the nature of large language models: A caution against anthropocentrism
por: Speed, Ann
Publicado: (2023)
por: Speed, Ann
Publicado: (2023)
The opportunities and risks of large language models in mental health
por: Lawrence, Hannah R., et al.
Publicado: (2024)
por: Lawrence, Hannah R., et al.
Publicado: (2024)
Large language models provide unsafe answers to patient-posed medical questions
por: Draelos, Rachel L., et al.
Publicado: (2025)
por: Draelos, Rachel L., et al.
Publicado: (2025)
A validity-guided workflow for robust large language model research in psychology
por: Lin, Zhicheng
Publicado: (2025)
por: Lin, Zhicheng
Publicado: (2025)
Evidence of a log scaling law for political persuasion with large language models
por: Hackenburg, Kobi, et al.
Publicado: (2024)
por: Hackenburg, Kobi, et al.
Publicado: (2024)
Humans overrely on overconfident language models, across languages
por: Rathi, Neil, et al.
Publicado: (2025)
por: Rathi, Neil, et al.
Publicado: (2025)
Recourse for reclamation: Chatting with generative language models
por: Chien, Jennifer, et al.
Publicado: (2024)
por: Chien, Jennifer, et al.
Publicado: (2024)
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
por: Seßler, Kathrin, et al.
Publicado: (2024)
por: Seßler, Kathrin, et al.
Publicado: (2024)
Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis
por: Armitage, Richard
Publicado: (2025)
por: Armitage, Richard
Publicado: (2025)
BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
por: Hang, Chi, et al.
Publicado: (2025)
por: Hang, Chi, et al.
Publicado: (2025)
Revisiting Active Learning under (Human) Label Variation
por: Gruber, Cornelia, et al.
Publicado: (2025)
por: Gruber, Cornelia, et al.
Publicado: (2025)
Fabricating Paper Circuits with Subtractive Processing
por: Yang, Ruhan, et al.
Publicado: (2024)
por: Yang, Ruhan, et al.
Publicado: (2024)
PhDGPT: Introducing a psychometric and linguistic dataset about how large language models perceive graduate students and professors in psychology
por: De Duro, Edoardo Sebastiano, et al.
Publicado: (2024)
por: De Duro, Edoardo Sebastiano, et al.
Publicado: (2024)
Large Language Models Predict Functional Outcomes after Acute Ischemic Stroke
por: Kapoor, Anjali K., et al.
Publicado: (2026)
por: Kapoor, Anjali K., et al.
Publicado: (2026)
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
por: Das, Dipto, et al.
Publicado: (2025)
por: Das, Dipto, et al.
Publicado: (2025)
A technology-oriented mapping of the language and translation industry: Analysing stakeholder values and their potential implication for translation pedagogy
por: Ginel, María Isabel Rivas, et al.
Publicado: (2026)
por: Ginel, María Isabel Rivas, et al.
Publicado: (2026)
When Avatars Have Personality: Effects on Engagement and Communication in Immersive Medical Training
por: Dollis, Julia S., et al.
Publicado: (2025)
por: Dollis, Julia S., et al.
Publicado: (2025)
VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs
por: Xiang, Yurui, et al.
Publicado: (2026)
por: Xiang, Yurui, et al.
Publicado: (2026)
Alzheimer's disease detection based on large language model prompt engineering
por: Zheng, Tian, et al.
Publicado: (2025)
por: Zheng, Tian, et al.
Publicado: (2025)
TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
por: Zhong, Shu, et al.
Publicado: (2024)
por: Zhong, Shu, et al.
Publicado: (2024)
Vibe Coding, Interface Flattening
por: Jin, Hongrui
Publicado: (2025)
por: Jin, Hongrui
Publicado: (2025)
Users Mispredict Their Own Preferences for AI Writing Assistance
por: Lai, Vivian, et al.
Publicado: (2026)
por: Lai, Vivian, et al.
Publicado: (2026)
Towards a copilot in BIM authoring tool using a large language model-based agent for intelligent human-machine interaction
por: Du, Changyu, et al.
Publicado: (2024)
por: Du, Changyu, et al.
Publicado: (2024)
Using a Human-AI Teaming Approach to Create and Curate Scientific Datasets with the SCILIRE System
por: Bölücü, Necva, et al.
Publicado: (2026)
por: Bölücü, Necva, et al.
Publicado: (2026)
Artificial intelligence language technologies in multilingual healthcare: Grand challenges ahead
por: Briva-Iglesias, Vicent
Publicado: (2026)
por: Briva-Iglesias, Vicent
Publicado: (2026)
The production of meaning in the processing of natural language
por: Agostino, Christopher J., et al.
Publicado: (2026)
por: Agostino, Christopher J., et al.
Publicado: (2026)
Towards a cognitive architecture to enable natural language interaction in co-constructive task learning
por: Scheibl, Manuel, et al.
Publicado: (2025)
por: Scheibl, Manuel, et al.
Publicado: (2025)
Ejemplares similares
-
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
por: Vishwanath, Krithik, et al.
Publicado: (2025) -
MedMobile: A mobile-sized language model with clinical capabilities
por: Vishwanath, Krithik, et al.
Publicado: (2024) -
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
por: Vishwanath, Krithik, et al.
Publicado: (2025) -
CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications
por: Alyakin, Anton, et al.
Publicado: (2025) -
Medication counseling with large language models: balancing flexibility and rigidity
por: Sabel, Joar, et al.
Publicado: (2025)