Salvato in:
| Autori principali: | V., Diego Fajardo, Proniakin, Oleksii, Gruber, Victoria-Elisabeth, Marinescu, Razvan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2601.04195 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automatic Replication of LLM Mistakes in Medical Conversations
di: Proniakin, Oleksii, et al.
Pubblicazione: (2025)
di: Proniakin, Oleksii, et al.
Pubblicazione: (2025)
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
di: Ahmed, Md Shamim, et al.
Pubblicazione: (2026)
di: Ahmed, Md Shamim, et al.
Pubblicazione: (2026)
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026)
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
di: Anyaegbuna, Chukwuebuka, et al.
Pubblicazione: (2026)
di: Anyaegbuna, Chukwuebuka, et al.
Pubblicazione: (2026)
Towards Leveraging Large Language Models for Automated Medical Q&A Evaluation
di: Krolik, Jack, et al.
Pubblicazione: (2024)
di: Krolik, Jack, et al.
Pubblicazione: (2024)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
di: Mehenni, Gaya, et al.
Pubblicazione: (2025)
di: Mehenni, Gaya, et al.
Pubblicazione: (2025)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
di: Manczak, Blazej, et al.
Pubblicazione: (2025)
di: Manczak, Blazej, et al.
Pubblicazione: (2025)
Evaluating the Challenges of LLMs in Real-world Medical Follow-up: A Comparative Study and An Optimized Framework
di: Liu, Jinyan, et al.
Pubblicazione: (2025)
di: Liu, Jinyan, et al.
Pubblicazione: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
di: Shah, Aaryan, et al.
Pubblicazione: (2026)
di: Shah, Aaryan, et al.
Pubblicazione: (2026)
How to Evaluate Medical AI
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
di: Zhang, Yiqing, et al.
Pubblicazione: (2026)
di: Zhang, Yiqing, et al.
Pubblicazione: (2026)
ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning
di: Ahmed, Md Shamim, et al.
Pubblicazione: (2026)
di: Ahmed, Md Shamim, et al.
Pubblicazione: (2026)
A Method for the Architecture of a Medical Vertical Large Language Model Based on Deepseek R1
di: Zhang, Mingda, et al.
Pubblicazione: (2025)
di: Zhang, Mingda, et al.
Pubblicazione: (2025)
CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification
di: Ye, Severin, et al.
Pubblicazione: (2026)
di: Ye, Severin, et al.
Pubblicazione: (2026)
Continuous Predictive Modeling of Clinical Notes and ICD Codes in Patient Health Records
di: Caralt, Mireia Hernandez, et al.
Pubblicazione: (2024)
di: Caralt, Mireia Hernandez, et al.
Pubblicazione: (2024)
The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
di: Alba, Charles, et al.
Pubblicazione: (2024)
di: Alba, Charles, et al.
Pubblicazione: (2024)
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
di: Petrova, Tatiana, et al.
Pubblicazione: (2026)
di: Petrova, Tatiana, et al.
Pubblicazione: (2026)
Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction
di: Ketir, Si-Belkacem Yamine, et al.
Pubblicazione: (2026)
di: Ketir, Si-Belkacem Yamine, et al.
Pubblicazione: (2026)
What distinguishes conspiracy from critical narratives? A computational analysis of oppositional discourse
di: Korenčić, Damir, et al.
Pubblicazione: (2024)
di: Korenčić, Damir, et al.
Pubblicazione: (2024)
Using Letter Positional Probabilities to Assess Word Complexity
di: Dalvean, Michael
Pubblicazione: (2024)
di: Dalvean, Michael
Pubblicazione: (2024)
Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis
di: Purushothama, Abhishek, et al.
Pubblicazione: (2025)
di: Purushothama, Abhishek, et al.
Pubblicazione: (2025)
BenCSSmark: Making the Social Sciences Count in LLM Research
di: Chatelain, Arnault, et al.
Pubblicazione: (2026)
di: Chatelain, Arnault, et al.
Pubblicazione: (2026)
LLMs Generate Kitsch
di: Klinge, Xenia, et al.
Pubblicazione: (2026)
di: Klinge, Xenia, et al.
Pubblicazione: (2026)
Integrating clinical reasoning into large language model-based diagnosis through etiology-aware attention steering
di: Li, Peixian, et al.
Pubblicazione: (2025)
di: Li, Peixian, et al.
Pubblicazione: (2025)
The Representational Alignment between Humans and Language Models is implicitly driven by a Concreteness Effect
di: Iaia, Cosimo, et al.
Pubblicazione: (2025)
di: Iaia, Cosimo, et al.
Pubblicazione: (2025)
Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
di: Castro-Gonzalez, Leonardo, et al.
Pubblicazione: (2024)
di: Castro-Gonzalez, Leonardo, et al.
Pubblicazione: (2024)
Comparative Study of Large Language Models on Chinese Film Script Continuation: An Empirical Analysis Based on GPT-5.2 and Qwen-Max
di: Cao, Yuxuan, et al.
Pubblicazione: (2026)
di: Cao, Yuxuan, et al.
Pubblicazione: (2026)
Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech
di: Floresca, Rez Samantha Z., et al.
Pubblicazione: (2026)
di: Floresca, Rez Samantha Z., et al.
Pubblicazione: (2026)
Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in Lebanon
di: Awwad, Ghadeer, et al.
Pubblicazione: (2025)
di: Awwad, Ghadeer, et al.
Pubblicazione: (2025)
Where is my Glass Slipper? AI, Poetry and Art
di: Pagiaslis, Anastasios P.
Pubblicazione: (2025)
di: Pagiaslis, Anastasios P.
Pubblicazione: (2025)
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
di: Elganayni, Mohamed Hesham, et al.
Pubblicazione: (2026)
di: Elganayni, Mohamed Hesham, et al.
Pubblicazione: (2026)
Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
di: Ge, Zhuohan, et al.
Pubblicazione: (2025)
di: Ge, Zhuohan, et al.
Pubblicazione: (2025)
AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
di: Linzmayer, Robin, et al.
Pubblicazione: (2026)
di: Linzmayer, Robin, et al.
Pubblicazione: (2026)
Reasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
di: Shih, Yu-Fei, et al.
Pubblicazione: (2025)
di: Shih, Yu-Fei, et al.
Pubblicazione: (2025)
BLT: Can Large Language Models Handle Basic Legal Text?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models
di: Cuscito, Miriam, et al.
Pubblicazione: (2024)
di: Cuscito, Miriam, et al.
Pubblicazione: (2024)
Chronic pain patient narratives allow for the estimation of current pain intensity
di: Nunes, Diogo A. P., et al.
Pubblicazione: (2022)
di: Nunes, Diogo A. P., et al.
Pubblicazione: (2022)
Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding
di: Rollman, John C., et al.
Pubblicazione: (2025)
di: Rollman, John C., et al.
Pubblicazione: (2025)
MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering
di: Ness, Robert Osazuwa, et al.
Pubblicazione: (2024)
di: Ness, Robert Osazuwa, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Automatic Replication of LLM Mistakes in Medical Conversations
di: Proniakin, Oleksii, et al.
Pubblicazione: (2025) -
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
di: Ahmed, Md Shamim, et al.
Pubblicazione: (2026) -
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
di: Kiet, Huynh Trung, et al.
Pubblicazione: (2026) -
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
di: Anyaegbuna, Chukwuebuka, et al.
Pubblicazione: (2026) -
Towards Leveraging Large Language Models for Automated Medical Q&A Evaluation
di: Krolik, Jack, et al.
Pubblicazione: (2024)