Saved in:
| Main Authors: | Shen, Yuan, Wu, Xiaojun, Yu, Linghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.11544 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ELMTEX: Fine-Tuning Large Language Models for Structured Clinical Information Extraction. A Case Study on Clinical Reports
by: Guluzade, Aynur, et al.
Published: (2025)
by: Guluzade, Aynur, et al.
Published: (2025)
Evaluating Large Language Models for IUCN Red List Species Information
by: Uryu, Shinya
Published: (2025)
by: Uryu, Shinya
Published: (2025)
End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians
by: Shah, Aaryan, et al.
Published: (2026)
by: Shah, Aaryan, et al.
Published: (2026)
A Method for the Architecture of a Medical Vertical Large Language Model Based on Deepseek R1
by: Zhang, Mingda, et al.
Published: (2025)
by: Zhang, Mingda, et al.
Published: (2025)
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
by: Shah, Aaryan, et al.
Published: (2026)
by: Shah, Aaryan, et al.
Published: (2026)
GPTON: Generative Pre-trained Transformers enhanced with Ontology Narration for accurate annotation of biological data
by: Li, Rongbin, et al.
Published: (2024)
by: Li, Rongbin, et al.
Published: (2024)
Ask WhAI:Probing Belief Formation in Role-Primed LLM Agents
by: Moore, Keith, et al.
Published: (2025)
by: Moore, Keith, et al.
Published: (2025)
Interpretability without actionability: mechanistic methods cannot correct language model errors despite near-perfect internal representations
by: Basu, Sanjay, et al.
Published: (2026)
by: Basu, Sanjay, et al.
Published: (2026)
AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
by: Linzmayer, Robin, et al.
Published: (2026)
by: Linzmayer, Robin, et al.
Published: (2026)
Igea: a Decoder-Only Language Model for Biomedical Text Generation in Italian
by: Buonocore, Tommaso Mario, et al.
Published: (2024)
by: Buonocore, Tommaso Mario, et al.
Published: (2024)
CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening
by: Lorenzoni, Giuliano, et al.
Published: (2026)
by: Lorenzoni, Giuliano, et al.
Published: (2026)
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
by: V., Diego Fajardo, et al.
Published: (2025)
by: V., Diego Fajardo, et al.
Published: (2025)
Model selection meets clinical semantics: Optimizing ICD-10-CM prediction via LLM-as-Judge evaluation, redundancy-aware sampling, and section-aware fine-tuning
by: Dai, Hong-Jie, et al.
Published: (2025)
by: Dai, Hong-Jie, et al.
Published: (2025)
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
by: Sakhovskiy, Andrey, et al.
Published: (2025)
by: Sakhovskiy, Andrey, et al.
Published: (2025)
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
BioAlchemy: Distilling Biological Literature into Reasoning-Ready Reinforcement Learning Training Data
by: Hsu, Brian, et al.
Published: (2026)
by: Hsu, Brian, et al.
Published: (2026)
OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
by: Tao, Xinli, et al.
Published: (2025)
by: Tao, Xinli, et al.
Published: (2025)
Fine-Tuning Open-Weight Language Models to Deliver Cognitive Behavioral Therapy for Depression: A Feasibility Study
by: Tahir, Talha
Published: (2024)
by: Tahir, Talha
Published: (2024)
The use of GPT-4o and Other Large Language Models for the Improvement and Design of Self-Assessment Scales for Measurement of Interpersonal Communication Skills
by: Bubaš, Goran
Published: (2024)
by: Bubaš, Goran
Published: (2024)
Curated AI beats frontier LLMs at pharma asset discovery
by: Kidziński, Łukasz, et al.
Published: (2026)
by: Kidziński, Łukasz, et al.
Published: (2026)
Classifiers of Data Sharing Statements in Clinical Trial Records
by: Mamaghani, Saber Jelodari, et al.
Published: (2025)
by: Mamaghani, Saber Jelodari, et al.
Published: (2025)
CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification
by: Ye, Severin, et al.
Published: (2026)
by: Ye, Severin, et al.
Published: (2026)
On Fusing ChatGPT and Ensemble Learning in Discon-tinuous Named Entity Recognition in Health Corpora
by: Chen, Tzu-Chieh, et al.
Published: (2024)
by: Chen, Tzu-Chieh, et al.
Published: (2024)
Evaluating the Challenges of LLMs in Real-world Medical Follow-up: A Comparative Study and An Optimized Framework
by: Liu, Jinyan, et al.
Published: (2025)
by: Liu, Jinyan, et al.
Published: (2025)
Expertise Is What We Want
by: Ashworth, Alan, et al.
Published: (2025)
by: Ashworth, Alan, et al.
Published: (2025)
Reshaping Free-Text Radiology Notes Into Structured Reports With Generative Transformers
by: Bergomi, Laura, et al.
Published: (2024)
by: Bergomi, Laura, et al.
Published: (2024)
Advancing Italian Biomedical Information Extraction with Transformers-based Models: Methodological Insights and Multicenter Practical Application
by: Crema, Claudio, et al.
Published: (2023)
by: Crema, Claudio, et al.
Published: (2023)
The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
by: Alba, Charles, et al.
Published: (2024)
by: Alba, Charles, et al.
Published: (2024)
ARGUS: Seeing the Influence of Narrative Features on Persuasion in Argumentative Texts
by: Nabhani, Sara, et al.
Published: (2026)
by: Nabhani, Sara, et al.
Published: (2026)
DrugReasoner: Interpretable Drug Approval Prediction with a Reasoning-augmented Language Model
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2025)
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2025)
Temporal Relation Extraction in Clinical Texts: A Span-based Graph Transformer Approach
by: Chaturvedi, Rochana, et al.
Published: (2025)
by: Chaturvedi, Rochana, et al.
Published: (2025)
DreamNet: A Multimodal Framework for Semantic and Emotional Analysis of Sleep Narratives
by: Panchagnula, Tapasvi
Published: (2025)
by: Panchagnula, Tapasvi
Published: (2025)
Performance of Large Language Models in Supporting Medical Diagnosis and Treatment
by: Sousa, Diogo, et al.
Published: (2025)
by: Sousa, Diogo, et al.
Published: (2025)
MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
by: Li, Yingyun, et al.
Published: (2026)
by: Li, Yingyun, et al.
Published: (2026)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
by: Manczak, Blazej, et al.
Published: (2025)
by: Manczak, Blazej, et al.
Published: (2025)
DALL-M: Context-Aware Clinical Data Augmentation with LLMs
by: Hsieh, Chihcheng, et al.
Published: (2024)
by: Hsieh, Chihcheng, et al.
Published: (2024)
The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games
by: Mozikov, Mikhail, et al.
Published: (2024)
by: Mozikov, Mikhail, et al.
Published: (2024)
From RAG to QA-RAG: Integrating Generative AI for Pharmaceutical Regulatory Compliance Process
by: Kim, Jaewoong, et al.
Published: (2024)
by: Kim, Jaewoong, et al.
Published: (2024)
Contrastive learning of T cell receptor representations
by: Nagano, Yuta, et al.
Published: (2024)
by: Nagano, Yuta, et al.
Published: (2024)
PerkwE_COQA: Enhanced Persian Conversational Question Answering by combining contextual keyword extraction with Large Language Models
by: Moradbeiki, Pardis, et al.
Published: (2024)
by: Moradbeiki, Pardis, et al.
Published: (2024)
Similar Items
-
ELMTEX: Fine-Tuning Large Language Models for Structured Clinical Information Extraction. A Case Study on Clinical Reports
by: Guluzade, Aynur, et al.
Published: (2025) -
Evaluating Large Language Models for IUCN Red List Species Information
by: Uryu, Shinya
Published: (2025) -
End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians
by: Shah, Aaryan, et al.
Published: (2026) -
A Method for the Architecture of a Medical Vertical Large Language Model Based on Deepseek R1
by: Zhang, Mingda, et al.
Published: (2025) -
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
by: Shah, Aaryan, et al.
Published: (2026)