Synthetic Clinical Notes for Rare ICD Codes: A Data-Centric Framework for Long-Tail Medical Coding
Fuente:
arXiv
Saved in:
| Main Authors: | Vo, Truong, Wu, Weiyi, Ding, Kaize |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding
by: Chen, Tianlei, et al.
Published: (2025)
by: Chen, Tianlei, et al.
Published: (2025)
A General Knowledge Injection Framework for ICD Coding
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Bridging the Version Gap: Multi-version Training Improves ICD Code Prediction, Especially for Rare Codes
by: Liu, Jinghui, et al.
Published: (2026)
by: Liu, Jinghui, et al.
Published: (2026)
Training a Large Language Model for Medical Coding Using Privacy-Preserving Synthetic Clinical Data
by: Cook, John, et al.
Published: (2026)
by: Cook, John, et al.
Published: (2026)
From Documents to Spans: Scalable Supervision for Evidence-Based ICD Coding with LLMs
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
MKE-Coder: Multi-Axial Knowledge with Evidence Verification in ICD Coding for Chinese EMRs
by: You, Xinxin, et al.
Published: (2025)
by: You, Xinxin, et al.
Published: (2025)
RuCCoD: Towards Automated ICD Coding in Russian
by: Nesterov, Aleksandr, et al.
Published: (2025)
by: Nesterov, Aleksandr, et al.
Published: (2025)
CoRelation: Boosting Automatic ICD Coding Through Contextualized Code Relation Learning
by: Luo, Junyu, et al.
Published: (2024)
by: Luo, Junyu, et al.
Published: (2024)
Structured Information Matters: Explainable ICD Coding with Patient-Level Knowledge Graphs
by: Li, Mingyang, et al.
Published: (2025)
by: Li, Mingyang, et al.
Published: (2025)
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
TraceCoder: Towards Traceable ICD Coding via Multi-Source Knowledge Integration
by: Ren, Mucheng, et al.
Published: (2025)
by: Ren, Mucheng, et al.
Published: (2025)
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings
by: Gomes, Gonçalo, et al.
Published: (2024)
by: Gomes, Gonçalo, et al.
Published: (2024)
Coding-Free and Privacy-Preserving Agentic Framework for Data-Driven Clinical Research
by: Kim, Taehun, et al.
Published: (2026)
by: Kim, Taehun, et al.
Published: (2026)
Empowering Large Language Models for Textual Data Augmentation
by: Li, Yichuan, et al.
Published: (2024)
by: Li, Yichuan, et al.
Published: (2024)
MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs
by: Mianroodi, Ahmad Rezaie, et al.
Published: (2025)
by: Mianroodi, Ahmad Rezaie, et al.
Published: (2025)
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
by: Liu, Jinghui, et al.
Published: (2026)
by: Liu, Jinghui, et al.
Published: (2026)
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes
by: Kweon, Sunjun, et al.
Published: (2023)
by: Kweon, Sunjun, et al.
Published: (2023)
Unlocking Public Catalogues: Instruction-Tuning LLMs for ICD Coding of German Tumor Diagnoses
by: Lenz, Stefan, et al.
Published: (2025)
by: Lenz, Stefan, et al.
Published: (2025)
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
by: Xu, Ruiyao, et al.
Published: (2026)
by: Xu, Ruiyao, et al.
Published: (2026)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
by: Hua, Tianyu, et al.
Published: (2025)
by: Hua, Tianyu, et al.
Published: (2025)
LongCodeBench: Evaluating Coding LLMs at 1M Context Windows
by: Rando, Stefano, et al.
Published: (2025)
by: Rando, Stefano, et al.
Published: (2025)
Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning
by: Wu, John, et al.
Published: (2024)
by: Wu, John, et al.
Published: (2024)
MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
A Comparative Study on Automatic Coding of Medical Letters with Explainability
by: Glen, Jamie, et al.
Published: (2024)
by: Glen, Jamie, et al.
Published: (2024)
DeCode: Decoupling Content and Delivery for Medical QA
by: Ko, Po-Jen, et al.
Published: (2026)
by: Ko, Po-Jen, et al.
Published: (2026)
LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
by: Zhu, Liya, et al.
Published: (2025)
by: Zhu, Liya, et al.
Published: (2025)
Coding Agents are Effective Long-Context Processors
by: Cao, Weili, et al.
Published: (2026)
by: Cao, Weili, et al.
Published: (2026)
Multimodal Medical Code Tokenizer
by: Su, Xiaorui, et al.
Published: (2025)
by: Su, Xiaorui, et al.
Published: (2025)
Training Versatile Coding Agents in Synthetic Environments
by: Zhu, Yiqi, et al.
Published: (2025)
by: Zhu, Yiqi, et al.
Published: (2025)
In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided Search
by: Li, Huihan, et al.
Published: (2023)
by: Li, Huihan, et al.
Published: (2023)
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
by: Huang, Yiming, et al.
Published: (2024)
by: Huang, Yiming, et al.
Published: (2024)
Generating High Quality Synthetic Data for Dutch Medical Conversations
by: Kuan, Cecilia, et al.
Published: (2026)
by: Kuan, Cecilia, et al.
Published: (2026)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
by: Abacha, Asma Ben, et al.
Published: (2024)
by: Abacha, Asma Ben, et al.
Published: (2024)
ITERTL: An Iterative Framework for Fine-tuning LLMs for RTL Code Generation
by: Wu, Peiyang, et al.
Published: (2024)
by: Wu, Peiyang, et al.
Published: (2024)
MIRA: A Bilingual Benchmark for Medical Information Response Audit
by: Xu, Mengyu, et al.
Published: (2026)
by: Xu, Mengyu, et al.
Published: (2026)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
by: Anh, Tran Nguyen, et al.
Published: (2025)
by: Anh, Tran Nguyen, et al.
Published: (2025)
SciCode: A Research Coding Benchmark Curated by Scientists
by: Tian, Minyang, et al.
Published: (2024)
by: Tian, Minyang, et al.
Published: (2024)
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Similar Items
-
Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding
by: Chen, Tianlei, et al.
Published: (2025) -
A General Knowledge Injection Framework for ICD Coding
by: Zhang, Xu, et al.
Published: (2025) -
Bridging the Version Gap: Multi-version Training Improves ICD Code Prediction, Especially for Rare Codes
by: Liu, Jinghui, et al.
Published: (2026) -
Training a Large Language Model for Medical Coding Using Privacy-Preserving Synthetic Clinical Data
by: Cook, John, et al.
Published: (2026) -
From Documents to Spans: Scalable Supervision for Evidence-Based ICD Coding with LLMs
by: Zhang, Xu, et al.
Published: (2026)