SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
Fuente:
arXiv
Guardado en:
| Autores principales: | Posada, Jose D., Love, David, Datta, Somalee, Desai, Priya |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
por: Miranda, Michele, et al.
Publicado: (2026)
por: Miranda, Michele, et al.
Publicado: (2026)
Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks
por: Joshi, Ratnesh Kumar, et al.
Publicado: (2024)
por: Joshi, Ratnesh Kumar, et al.
Publicado: (2024)
Distilling Mathematical Reasoning Capabilities into Small Language Models
por: Zhu, Xunyu, et al.
Publicado: (2024)
por: Zhu, Xunyu, et al.
Publicado: (2024)
LLMs-in-the-Loop Part 2: Expert Small AI Models for Anonymization and De-identification of PHI Across Multiple Languages
por: Gunay, Murat, et al.
Publicado: (2024)
por: Gunay, Murat, et al.
Publicado: (2024)
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes
por: Kweon, Sunjun, et al.
Publicado: (2023)
por: Kweon, Sunjun, et al.
Publicado: (2023)
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
por: Liu, Jinghui, et al.
Publicado: (2026)
por: Liu, Jinghui, et al.
Publicado: (2026)
The Evolution of RWKV: Advancements in Efficient Language Modeling
por: Datta, Akul
Publicado: (2024)
por: Datta, Akul
Publicado: (2024)
Capturing Nuanced Preferences: Preference-Aligned Distillation for Small Language Models
por: Gu, Yanggan, et al.
Publicado: (2025)
por: Gu, Yanggan, et al.
Publicado: (2025)
SOD: Step-wise On-policy Distillation for Small Language Model Agents
por: Zhong, Qiyong, et al.
Publicado: (2026)
por: Zhong, Qiyong, et al.
Publicado: (2026)
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
por: Kang, Yue, et al.
Publicado: (2026)
por: Kang, Yue, et al.
Publicado: (2026)
Kakugo: Distillation of Low-Resource Languages into Small Language Models
por: Devine, Peter, et al.
Publicado: (2026)
por: Devine, Peter, et al.
Publicado: (2026)
Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation
por: Zhu, Xunyu, et al.
Publicado: (2024)
por: Zhu, Xunyu, et al.
Publicado: (2024)
Efficiency at Scale: Investigating the Performance of Diminutive Language Models in Clinical Tasks
por: Taylor, Niall, et al.
Publicado: (2024)
por: Taylor, Niall, et al.
Publicado: (2024)
DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models
por: Wang, Bowen, et al.
Publicado: (2024)
por: Wang, Bowen, et al.
Publicado: (2024)
Declarative Knowledge Distillation from Large Language Models for Visual Question Answering Datasets
por: Eiter, Thomas, et al.
Publicado: (2024)
por: Eiter, Thomas, et al.
Publicado: (2024)
Large Language Models as Universal Predictors? An Empirical Study on Small Tabular Datasets
por: Pavlidis, Nikolaos, et al.
Publicado: (2025)
por: Pavlidis, Nikolaos, et al.
Publicado: (2025)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
por: Liu, Xiaoze, et al.
Publicado: (2024)
por: Liu, Xiaoze, et al.
Publicado: (2024)
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
por: Song, Feifan, et al.
Publicado: (2024)
por: Song, Feifan, et al.
Publicado: (2024)
We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism
por: Priya, Priyanshu, et al.
Publicado: (2025)
por: Priya, Priyanshu, et al.
Publicado: (2025)
Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction
por: Chuang, Yao-Shun, et al.
Publicado: (2026)
por: Chuang, Yao-Shun, et al.
Publicado: (2026)
Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation
por: Cheng, Chuanqi, et al.
Publicado: (2025)
por: Cheng, Chuanqi, et al.
Publicado: (2025)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
por: Kang, Minki, et al.
Publicado: (2025)
por: Kang, Minki, et al.
Publicado: (2025)
SLaDe: A Portable Small Language Model Decompiler for Optimized Assembly
por: Armengol-Estapé, Jordi, et al.
Publicado: (2023)
por: Armengol-Estapé, Jordi, et al.
Publicado: (2023)
Chunk-Distilled Language Modeling
por: Li, Yanhong, et al.
Publicado: (2024)
por: Li, Yanhong, et al.
Publicado: (2024)
Extraction of Sleep Information from Clinical Notes of Patients with Alzheimer's Disease Using Natural Language Processing
por: Sivarajkumar, Sonish, et al.
Publicado: (2022)
por: Sivarajkumar, Sonish, et al.
Publicado: (2022)
Efficient Standardization of Clinical Notes using Large Language Models
por: Hier, Daniel B., et al.
Publicado: (2024)
por: Hier, Daniel B., et al.
Publicado: (2024)
GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions
por: Salierno, Giulio, et al.
Publicado: (2024)
por: Salierno, Giulio, et al.
Publicado: (2024)
Assessing the Quality of AI-Generated Clinical Notes: A Validated Evaluation of a Large Language Model Scribe
por: Palm, Erin, et al.
Publicado: (2025)
por: Palm, Erin, et al.
Publicado: (2025)
Importance of Prompt Optimisation for Error Detection in Medical Notes Using Language Models
por: Myles, Craig, et al.
Publicado: (2026)
por: Myles, Craig, et al.
Publicado: (2026)
Technical Report: Small Language Model for Japanese Clinical and Medicine
por: Watanabe, Shogo
Publicado: (2024)
por: Watanabe, Shogo
Publicado: (2024)
Domain-Adapted Small Language Models for Reliable Clinical Triage
por: Aljohani, Manar, et al.
Publicado: (2026)
por: Aljohani, Manar, et al.
Publicado: (2026)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
por: Trung, Bui The, et al.
Publicado: (2026)
por: Trung, Bui The, et al.
Publicado: (2026)
T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
por: Kang, Minki, et al.
Publicado: (2025)
por: Kang, Minki, et al.
Publicado: (2025)
Rethinking Scale: Deployment Trade-offs of Small Language Models under Agent Paradigms
por: Wang, Xinlin, et al.
Publicado: (2026)
por: Wang, Xinlin, et al.
Publicado: (2026)
Command A: An Enterprise-Ready Large Language Model
por: Cohere, Team, et al.
Publicado: (2025)
por: Cohere, Team, et al.
Publicado: (2025)
Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
por: Muneeb, Muhammad, et al.
Publicado: (2025)
por: Muneeb, Muhammad, et al.
Publicado: (2025)
Measuring Diversity in Synthetic Datasets
por: Zhu, Yuchang, et al.
Publicado: (2025)
por: Zhu, Yuchang, et al.
Publicado: (2025)
Unifying Ontology Construction and Semantic Alignment for Deterministic Enterprise Reasoning at Scale
por: Zhu, Hongyin
Publicado: (2026)
por: Zhu, Hongyin
Publicado: (2026)
Classification of Radiological Text in Small and Imbalanced Datasets in a Non-English Language
por: Beliveau, Vincent, et al.
Publicado: (2024)
por: Beliveau, Vincent, et al.
Publicado: (2024)
InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning
por: Xie, Congkai, et al.
Publicado: (2025)
por: Xie, Congkai, et al.
Publicado: (2025)
Ejemplares similares
-
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
por: Miranda, Michele, et al.
Publicado: (2026) -
Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks
por: Joshi, Ratnesh Kumar, et al.
Publicado: (2024) -
Distilling Mathematical Reasoning Capabilities into Small Language Models
por: Zhu, Xunyu, et al.
Publicado: (2024) -
LLMs-in-the-Loop Part 2: Expert Small AI Models for Anonymization and De-identification of PHI Across Multiple Languages
por: Gunay, Murat, et al.
Publicado: (2024) -
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes
por: Kweon, Sunjun, et al.
Publicado: (2023)