Following the Teacher's Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Cheng, Zhong, Chaoliang, Sun, Jun, Oishi, Yusuke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
von: Xia, Xiaojie, et al.
Veröffentlicht: (2026)
von: Xia, Xiaojie, et al.
Veröffentlicht: (2026)
FootstepNet: an Efficient Actor-Critic Method for Fast On-line Bipedal Footstep Planning and Forecasting
von: Gaspard, Clément, et al.
Veröffentlicht: (2024)
von: Gaspard, Clément, et al.
Veröffentlicht: (2024)
Automated Constraint Specification for Job Scheduling by Regulating Generative Model with Domain-Specific Representation
von: Shi, Yu-Zhe, et al.
Veröffentlicht: (2025)
von: Shi, Yu-Zhe, et al.
Veröffentlicht: (2025)
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
von: Kim, Seongho, et al.
Veröffentlicht: (2024)
von: Kim, Seongho, et al.
Veröffentlicht: (2024)
Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
von: Chen, Tianlei, et al.
Veröffentlicht: (2026)
von: Chen, Tianlei, et al.
Veröffentlicht: (2026)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
von: Fang, Luyang, et al.
Veröffentlicht: (2026)
von: Fang, Luyang, et al.
Veröffentlicht: (2026)
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
von: Zhong, Yisheng, et al.
Veröffentlicht: (2026)
von: Zhong, Yisheng, et al.
Veröffentlicht: (2026)
LLMs to Support a Domain Specific Knowledge Assistant
von: Lovin, Maria-Flavia
Veröffentlicht: (2025)
von: Lovin, Maria-Flavia
Veröffentlicht: (2025)
LLMs can Schedule
von: Abgaryan, Henrik, et al.
Veröffentlicht: (2024)
von: Abgaryan, Henrik, et al.
Veröffentlicht: (2024)
DSG-KD: Knowledge Distillation from Domain-Specific to General Language Models
von: Cho, Sangyeon, et al.
Veröffentlicht: (2024)
von: Cho, Sangyeon, et al.
Veröffentlicht: (2024)
How good are LLMs at Retrieving Documents in a Specific Domain?
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2025)
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2025)
All-in-One Tuning and Structural Pruning for Domain-Specific LLMs
von: Lu, Lei, et al.
Veröffentlicht: (2024)
von: Lu, Lei, et al.
Veröffentlicht: (2024)
PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs
von: Liu, An, et al.
Veröffentlicht: (2024)
von: Liu, An, et al.
Veröffentlicht: (2024)
Neuro-Symbolic Verification on Instruction Following of LLMs
von: Su, Yiming, et al.
Veröffentlicht: (2026)
von: Su, Yiming, et al.
Veröffentlicht: (2026)
Reasoning Scaffolding: Distilling the Flow of Thought from LLMs
von: Wen, Xiangyu, et al.
Veröffentlicht: (2025)
von: Wen, Xiangyu, et al.
Veröffentlicht: (2025)
Tracing Footsteps of Similar Cities: Modeling Urban Economic Vitality with Dynamic Inter-City Graph Embeddings
von: Li, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Li, Xiaofeng, et al.
Veröffentlicht: (2025)
Agentic Adversarial QA for Improving Domain-Specific LLMs
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
Dynamic Temperature Scheduler for Knowledge Distillation
von: Islam, Sibgat Ul, et al.
Veröffentlicht: (2025)
von: Islam, Sibgat Ul, et al.
Veröffentlicht: (2025)
Domain-Specific Improvement on Psychotherapy Chatbot Using Assistant
von: Kang, Cheng, et al.
Veröffentlicht: (2024)
von: Kang, Cheng, et al.
Veröffentlicht: (2024)
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
von: Nguyen, Duy, et al.
Veröffentlicht: (2026)
von: Nguyen, Duy, et al.
Veröffentlicht: (2026)
Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge Distillation
von: Zhan, Yu-Liang, et al.
Veröffentlicht: (2024)
von: Zhan, Yu-Liang, et al.
Veröffentlicht: (2024)
ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development
von: Wan, Borui, et al.
Veröffentlicht: (2024)
von: Wan, Borui, et al.
Veröffentlicht: (2024)
From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs
von: Chen, Yingjian, et al.
Veröffentlicht: (2026)
von: Chen, Yingjian, et al.
Veröffentlicht: (2026)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
von: Hu, Lanxiang, et al.
Veröffentlicht: (2024)
von: Hu, Lanxiang, et al.
Veröffentlicht: (2024)
Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation
von: Yan, Xu, et al.
Veröffentlicht: (2026)
von: Yan, Xu, et al.
Veröffentlicht: (2026)
Fine-Tuned Language Models for Domain-Specific Summarization and Tagging
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
von: Ravishankara, Mayank
Veröffentlicht: (2026)
von: Ravishankara, Mayank
Veröffentlicht: (2026)
Harmonizing Multi-Objective LLM Unlearning via Unified Domain Representation and Bidirectional Logit Distillation
von: Zhong, Yisheng, et al.
Veröffentlicht: (2026)
von: Zhong, Yisheng, et al.
Veröffentlicht: (2026)
DRAK: Unlocking Molecular Insights with Domain-Specific Retrieval-Augmented Knowledge in LLMs
von: Liu, Jinzhe, et al.
Veröffentlicht: (2024)
von: Liu, Jinzhe, et al.
Veröffentlicht: (2024)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
von: Chen, Rubing, et al.
Veröffentlicht: (2025)
von: Chen, Rubing, et al.
Veröffentlicht: (2025)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
von: Qin, Zongyue, et al.
Veröffentlicht: (2025)
von: Qin, Zongyue, et al.
Veröffentlicht: (2025)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation
von: Liang, Lei, et al.
Veröffentlicht: (2024)
von: Liang, Lei, et al.
Veröffentlicht: (2024)
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
von: Wei, Tong, et al.
Veröffentlicht: (2025)
von: Wei, Tong, et al.
Veröffentlicht: (2025)
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
von: Xie, Yuhan, et al.
Veröffentlicht: (2026)
von: Xie, Yuhan, et al.
Veröffentlicht: (2026)
From Guidelines to Guarantees: A Graph-Based Evaluation Harness for Domain-Specific Evaluation of LLMs
von: Lundin, Jessica M., et al.
Veröffentlicht: (2025)
von: Lundin, Jessica M., et al.
Veröffentlicht: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
von: Tian, Yijun, et al.
Veröffentlicht: (2024)
von: Tian, Yijun, et al.
Veröffentlicht: (2024)
Domain Specific Specialization in Low-Resource Settings: The Efficacy of Offline Response-Based Knowledge Distillation in Large Language Models
von: Aslan, Erdem, et al.
Veröffentlicht: (2026)
von: Aslan, Erdem, et al.
Veröffentlicht: (2026)
TOM: An Open-Source Tongue Segmentation Method with Multi-Teacher Distillation and Task-Specific Data Augmentation
von: Xie, Jiacheng, et al.
Veröffentlicht: (2025)
von: Xie, Jiacheng, et al.
Veröffentlicht: (2025)
R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
von: Jain, Raj, et al.
Veröffentlicht: (2025)
von: Jain, Raj, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
von: Xia, Xiaojie, et al.
Veröffentlicht: (2026) -
FootstepNet: an Efficient Actor-Critic Method for Fast On-line Bipedal Footstep Planning and Forecasting
von: Gaspard, Clément, et al.
Veröffentlicht: (2024) -
Automated Constraint Specification for Job Scheduling by Regulating Generative Model with Domain-Specific Representation
von: Shi, Yu-Zhe, et al.
Veröffentlicht: (2025) -
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
von: Kim, Seongho, et al.
Veröffentlicht: (2024) -
Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
von: Chen, Tianlei, et al.
Veröffentlicht: (2026)