Knowledge Hierarchy Guided Biological-Medical Dataset Distillation for Domain LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Xunxin, Wang, Chengrui, Long, Qingqing, Zhou, Yuanchun, Xiao, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training
by: Xiao, Meng, et al.
Published: (2025)
by: Xiao, Meng, et al.
Published: (2025)
BioRAG: A RAG-LLM Framework for Biological Question Reasoning
by: Wang, Chengrui, et al.
Published: (2024)
by: Wang, Chengrui, et al.
Published: (2024)
scReader: Prompting Large Language Models to Interpret scRNA-seq Data
by: Li, Cong, et al.
Published: (2024)
by: Li, Cong, et al.
Published: (2024)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoning
by: Xing, Wang, et al.
Published: (2026)
by: Xing, Wang, et al.
Published: (2026)
GeneSUM: Large Language Model-based Gene Summary Extraction
by: Chen, Zhijian, et al.
Published: (2024)
by: Chen, Zhijian, et al.
Published: (2024)
Distilling Closed-Source LLM's Knowledge for Locally Stable and Economic Biomedical Entity Linking
by: Ai, Yihao, et al.
Published: (2025)
by: Ai, Yihao, et al.
Published: (2025)
MiniPLM: Knowledge Distillation for Pre-Training Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering
by: Chen, Haotian, et al.
Published: (2026)
by: Chen, Haotian, et al.
Published: (2026)
Performance-Guided LLM Knowledge Distillation for Efficient Text Classification at Scale
by: Di Palo, Flavio, et al.
Published: (2024)
by: Di Palo, Flavio, et al.
Published: (2024)
LLM-based Privacy Data Augmentation Guided by Knowledge Distillation with a Distribution Tutor for Medical Text Classification
by: Song, Yiping, et al.
Published: (2024)
by: Song, Yiping, et al.
Published: (2024)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
by: Xiao, Yilin, et al.
Published: (2025)
by: Xiao, Yilin, et al.
Published: (2025)
Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets
by: Pan, Jiqun, et al.
Published: (2025)
by: Pan, Jiqun, et al.
Published: (2025)
Knowledge Distillation with Training Wheels
by: Liu, Guanlin, et al.
Published: (2025)
by: Liu, Guanlin, et al.
Published: (2025)
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generation
by: Han, Chengrui, et al.
Published: (2026)
by: Han, Chengrui, et al.
Published: (2026)
Warmup-Distill: Bridge the Distribution Mismatch between Teacher and Student before Knowledge Distillation
by: Sun, Zengkui, et al.
Published: (2025)
by: Sun, Zengkui, et al.
Published: (2025)
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation
by: Liang, Zi, et al.
Published: (2024)
by: Liang, Zi, et al.
Published: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
by: Li, Weiyuan, et al.
Published: (2025)
by: Li, Weiyuan, et al.
Published: (2025)
LLM-Oriented Token-Adaptive Knowledge Distillation
by: Xie, Xurong, et al.
Published: (2025)
by: Xie, Xurong, et al.
Published: (2025)
Entropy-Guided Token Dropout: Training Autoregressive Language Models with Limited Domain Data
by: Wang, Jiapeng, et al.
Published: (2025)
by: Wang, Jiapeng, et al.
Published: (2025)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
by: Fang, Luyang, et al.
Published: (2025)
by: Fang, Luyang, et al.
Published: (2025)
Utilizing Local Hierarchy with Adversarial Training for Hierarchical Text Classification
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping
by: Chen, Yijie, et al.
Published: (2025)
by: Chen, Yijie, et al.
Published: (2025)
LLM-MRD: LLM-Guided Multi-View Reasoning Distillation for Fake News Detection
by: Zhou, Weilin, et al.
Published: (2026)
by: Zhou, Weilin, et al.
Published: (2026)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
by: Guo, Chuan, et al.
Published: (2026)
by: Guo, Chuan, et al.
Published: (2026)
Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond
by: Li, Qiyuan, et al.
Published: (2025)
by: Li, Qiyuan, et al.
Published: (2025)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
by: Shen, Yiyang, et al.
Published: (2026)
by: Shen, Yiyang, et al.
Published: (2026)
TTPA: Token-level Tool-use Preference Alignment Training Framework with Fine-grained Evaluation
by: Huang, Chengrui, et al.
Published: (2025)
by: Huang, Chengrui, et al.
Published: (2025)
SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models
by: Qin, Chuan, et al.
Published: (2025)
by: Qin, Chuan, et al.
Published: (2025)
Self-Evolution Knowledge Distillation for LLM-based Machine Translation
by: Song, Yuncheng, et al.
Published: (2024)
by: Song, Yuncheng, et al.
Published: (2024)
EpilepsyLLM: Domain-Specific Large Language Model Fine-tuned with Epilepsy Medical Knowledge
by: Zhao, Xuyang, et al.
Published: (2024)
by: Zhao, Xuyang, et al.
Published: (2024)
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
SciTopic: Enhancing Topic Discovery in Scientific Literature through Advanced LLM
by: Li, Pengjiang, et al.
Published: (2025)
by: Li, Pengjiang, et al.
Published: (2025)
GOLD: Generalized Knowledge Distillation via Out-of-Distribution-Guided Language Data Generation
by: Gholami, Mohsen, et al.
Published: (2024)
by: Gholami, Mohsen, et al.
Published: (2024)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
$\mathcal{X}$-KD: General Experiential Knowledge Distillation for Large Language Models
by: Cai, Yuang, et al.
Published: (2026)
by: Cai, Yuang, et al.
Published: (2026)
From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization
by: Cui, Chaoqun, et al.
Published: (2026)
by: Cui, Chaoqun, et al.
Published: (2026)
Similar Items
-
Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training
by: Xiao, Meng, et al.
Published: (2025) -
BioRAG: A RAG-LLM Framework for Biological Question Reasoning
by: Wang, Chengrui, et al.
Published: (2024) -
scReader: Prompting Large Language Models to Interpret scRNA-seq Data
by: Li, Cong, et al.
Published: (2024) -
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025) -
LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoning
by: Xing, Wang, et al.
Published: (2026)