Rethinking the Role of Temperature in Large Language Model Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Luong, Hoang-Chau, Chen, Lingwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Consistently Informative Soft-Label Temperature for Knowledge Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Enhancing Graph Neural Networks with Limited Labeled Data by Actively Distilling Knowledge from Large Language Models
by: Li, Quan, et al.
Published: (2024)
by: Li, Quan, et al.
Published: (2024)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
by: Li, Yaxuan, et al.
Published: (2026)
by: Li, Yaxuan, et al.
Published: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026)
by: Zhao, Hanyang, et al.
Published: (2026)
LoReC: Rethinking Large Language Models for Graph Data Analysis
by: Zhan, Hongyu, et al.
Published: (2026)
by: Zhan, Hongyu, et al.
Published: (2026)
Large Language Model Guided Knowledge Distillation for Time Series Anomaly Detection
by: Liu, Chen, et al.
Published: (2024)
by: Liu, Chen, et al.
Published: (2024)
Rethinking the Role of Proxy Rewards in Language Model Alignment
by: Kim, Sungdong, et al.
Published: (2024)
by: Kim, Sungdong, et al.
Published: (2024)
Leveraging Large Language Models for Suicide Detection on Social Media with Limited Labels
by: Nguyen, Vy, et al.
Published: (2024)
by: Nguyen, Vy, et al.
Published: (2024)
Curriculum Learning-Guided Progressive Distillation in Large Language Models
by: Cao, Jincheng, et al.
Published: (2026)
by: Cao, Jincheng, et al.
Published: (2026)
Distillation of Large Language Models via Concrete Score Matching
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
Rethinking Data Mixing from the Perspective of Large Language Models
by: Xu, Yuanjian, et al.
Published: (2026)
by: Xu, Yuanjian, et al.
Published: (2026)
Rethinking Interpretability in the Era of Large Language Models
by: Singh, Chandan, et al.
Published: (2024)
by: Singh, Chandan, et al.
Published: (2024)
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Rethinking Momentum Knowledge Distillation in Online Continual Learning
by: Michel, Nicolas, et al.
Published: (2023)
by: Michel, Nicolas, et al.
Published: (2023)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
Large Language Models Explore by Latent Distilling
by: Zeng, Yuanhao, et al.
Published: (2026)
by: Zeng, Yuanhao, et al.
Published: (2026)
Delta Knowledge Distillation for Large Language Models
by: Cao, Yihan, et al.
Published: (2025)
by: Cao, Yihan, et al.
Published: (2025)
Structured Agent Distillation for Large Language Model
by: Liu, Jun, et al.
Published: (2025)
by: Liu, Jun, et al.
Published: (2025)
Teach Harder, Learn Poorer: Rethinking Hard Sample Distillation for GNN-to-MLP Knowledge Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation
by: Song, Siqing, et al.
Published: (2026)
by: Song, Siqing, et al.
Published: (2026)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
by: Park, Seonghyeon, et al.
Published: (2026)
by: Park, Seonghyeon, et al.
Published: (2026)
Lillama: Large Language Models Compression via Low-Rank Feature Distillation
by: Sy, Yaya, et al.
Published: (2024)
by: Sy, Yaya, et al.
Published: (2024)
Large Language Model Distilling Medication Recommendation Model
by: Liu, Qidong, et al.
Published: (2024)
by: Liu, Qidong, et al.
Published: (2024)
Dynamic Temperature Scheduler for Knowledge Distillation
by: Islam, Sibgat Ul, et al.
Published: (2025)
by: Islam, Sibgat Ul, et al.
Published: (2025)
T-LLM: Teaching Large Language Models to Forecast Time Series via Temporal Distillation
by: Guo, Suhan, et al.
Published: (2026)
by: Guo, Suhan, et al.
Published: (2026)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
by: Nguyen, Nam V., et al.
Published: (2024)
by: Nguyen, Nam V., et al.
Published: (2024)
Privileged Information Distillation for Language Models
by: Penaloza, Emiliano, et al.
Published: (2026)
by: Penaloza, Emiliano, et al.
Published: (2026)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
by: Lee, Ivan, et al.
Published: (2025)
by: Lee, Ivan, et al.
Published: (2025)
Beyond Frequency: The Role of Redundancy in Large Language Model Memorization
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
NanoKnow: How to Know What Your Language Model Knows
by: Gu, Lingwei, et al.
Published: (2026)
by: Gu, Lingwei, et al.
Published: (2026)
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
by: Zhang, Shu-Hao, et al.
Published: (2026)
by: Zhang, Shu-Hao, et al.
Published: (2026)
LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
by: Zhu, Lingwei, et al.
Published: (2023)
by: Zhu, Lingwei, et al.
Published: (2023)
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
NNGPT: Rethinking AutoML with Large Language Models
by: Kochnev, Roman, et al.
Published: (2025)
by: Kochnev, Roman, et al.
Published: (2025)
Rethinking the Intermediate Features in Adversarial Attacks: Misleading Robotic Models via Adversarial Distillation
by: Zhao, Ke, et al.
Published: (2024)
by: Zhao, Ke, et al.
Published: (2024)
A Dual-Space Framework for General Knowledge Distillation of Large Language Models
by: Zhang, Xue, et al.
Published: (2025)
by: Zhang, Xue, et al.
Published: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
by: Jiang, Yuxian, et al.
Published: (2025)
by: Jiang, Yuxian, et al.
Published: (2025)
Similar Items
-
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026) -
Consistently Informative Soft-Label Temperature for Knowledge Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026) -
Enhancing Graph Neural Networks with Limited Labeled Data by Actively Distilling Knowledge from Large Language Models
by: Li, Quan, et al.
Published: (2024) -
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
by: Zimmer, Matthieu, et al.
Published: (2025) -
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
by: Li, Yaxuan, et al.
Published: (2026)