Consistently Informative Soft-Label Temperature for Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Luong, Hoang-Chau, Van Vo, Nghia, Zhao, Kaiqi, Chen, Lingwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking the Role of Temperature in Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Enhancing Graph Neural Networks with Limited Labeled Data by Actively Distilling Knowledge from Large Language Models
by: Li, Quan, et al.
Published: (2024)
by: Li, Quan, et al.
Published: (2024)
Self-Supervised Quantization-Aware Knowledge Distillation
by: Zhao, Kaiqi, et al.
Published: (2024)
by: Zhao, Kaiqi, et al.
Published: (2024)
RePCS: Diagnosing Data Memorization in LLM-Powered Retrieval-Augmented Generation
by: Anh, Le Vu, et al.
Published: (2025)
by: Anh, Le Vu, et al.
Published: (2025)
Improve Knowledge Distillation via Label Revision and Data Selection
by: Lan, Weichao, et al.
Published: (2024)
by: Lan, Weichao, et al.
Published: (2024)
Dynamic Temperature Scheduler for Knowledge Distillation
by: Islam, Sibgat Ul, et al.
Published: (2025)
by: Islam, Sibgat Ul, et al.
Published: (2025)
Multi-Label Knowledge Distillation
by: Yang, Penghui, et al.
Published: (2023)
by: Yang, Penghui, et al.
Published: (2023)
Instance Temperature Knowledge Distillation
by: Zhang, Zhengbo, et al.
Published: (2024)
by: Zhang, Zhengbo, et al.
Published: (2024)
Learning Surrogates for Offline Black-Box Optimization via Gradient Matching
by: Hoang, Minh, et al.
Published: (2025)
by: Hoang, Minh, et al.
Published: (2025)
Black-Box Optimization From Small Offline Datasets via Meta Learning with Synthetic Tasks
by: Fadhel, Azza, et al.
Published: (2026)
by: Fadhel, Azza, et al.
Published: (2026)
DANA: Domain-Aware Neurosymbolic Agents for Consistency and Accuracy
by: Luong, Vinh, et al.
Published: (2024)
by: Luong, Vinh, et al.
Published: (2024)
Soft Label Pruning and Quantization for Large-Scale Dataset Distillation
by: Lingao, Xiao, et al.
Published: (2026)
by: Lingao, Xiao, et al.
Published: (2026)
Dynamic Temperature Knowledge Distillation
by: Wei, Yukang, et al.
Published: (2024)
by: Wei, Yukang, et al.
Published: (2024)
Offline Model-Based Optimization via Policy-Guided Gradient Search
by: Chemingui, Yassine, et al.
Published: (2024)
by: Chemingui, Yassine, et al.
Published: (2024)
FedKDX: Federated Learning with Negative Knowledge Distillation for Enhanced Healthcare AI Systems
by: Pham, Quang-Tu, et al.
Published: (2026)
by: Pham, Quang-Tu, et al.
Published: (2026)
Halal or Not: Knowledge Graph Completion for Predicting Cultural Appropriateness of Daily Products
by: Hoang, Van Thuy, et al.
Published: (2025)
by: Hoang, Van Thuy, et al.
Published: (2025)
Efficient Bilevel Optimization for Meta Label Correction in Noisy Label Learning
by: Nguyen, Ba Hoang Anh, et al.
Published: (2026)
by: Nguyen, Ba Hoang Anh, et al.
Published: (2026)
On the Memorization of Consistency Distillation for Diffusion Models
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
Smooth-Distill: A Self-distillation Framework for Multitask Learning with Wearable Sensor Data
by: Vu, Hoang-Dieu, et al.
Published: (2025)
by: Vu, Hoang-Dieu, et al.
Published: (2025)
Soft-Label Training Preserves Epistemic Uncertainty
by: Singh, Agamdeep, et al.
Published: (2025)
by: Singh, Agamdeep, et al.
Published: (2025)
Pre-training Graph Neural Networks on Molecules by Using Subgraph-Conditioned Graph Information Bottleneck
by: Hoang, Van Thuy, et al.
Published: (2024)
by: Hoang, Van Thuy, et al.
Published: (2024)
Enhancing Distribution and Label Consistency for Graph Out-of-Distribution Generalization
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Probabilistic Federated Prompt-Tuning with Non-IID and Imbalanced Data
by: Weng, Pei-Yau, et al.
Published: (2025)
by: Weng, Pei-Yau, et al.
Published: (2025)
Pre-training Graph Neural Networks on 2D and 3D Molecular Structures by using Multi-View Conditional Information Bottleneck
by: Hoang, Van Thuy, et al.
Published: (2025)
by: Hoang, Van Thuy, et al.
Published: (2025)
Online Adversarial Knowledge Distillation for Graph Neural Networks
by: Wang, Can, et al.
Published: (2021)
by: Wang, Can, et al.
Published: (2021)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
by: Chen, Jinyin, et al.
Published: (2024)
by: Chen, Jinyin, et al.
Published: (2024)
A Note on Knowledge Distillation Loss Function for Object Classification
by: Chen, Defang
Published: (2021)
by: Chen, Defang
Published: (2021)
AdaGMLP: AdaBoosting GNN-to-MLP Knowledge Distillation
by: Lu, Weigang, et al.
Published: (2024)
by: Lu, Weigang, et al.
Published: (2024)
Consistency-Guided Temperature Scaling Using Style and Content Information for Out-of-Domain Calibration
by: Choi, Wonjeong, et al.
Published: (2024)
by: Choi, Wonjeong, et al.
Published: (2024)
Leveraging Large Language Models for Suicide Detection on Social Media with Limited Labels
by: Nguyen, Vy, et al.
Published: (2024)
by: Nguyen, Vy, et al.
Published: (2024)
R.I.P.: A Simple Black-box Attack on Continual Test-time Adaptation
by: Hoang, Trung-Hieu, et al.
Published: (2024)
by: Hoang, Trung-Hieu, et al.
Published: (2024)
Self-Supervised Representation Learning for Geospatial Objects: A Survey
by: Chen, Yile, et al.
Published: (2024)
by: Chen, Yile, et al.
Published: (2024)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
by: Zhu, Lingwei, et al.
Published: (2023)
by: Zhu, Lingwei, et al.
Published: (2023)
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
by: Wang, Wenshuo, et al.
Published: (2025)
by: Wang, Wenshuo, et al.
Published: (2025)
Graph Knowledge Distillation to Mixture of Experts
by: Rumiantsev, Pavel, et al.
Published: (2024)
by: Rumiantsev, Pavel, et al.
Published: (2024)
Membership and Memorization in LLM Knowledge Distillation
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
by: Vo, Thanh Vinh, et al.
Published: (2025)
by: Vo, Thanh Vinh, et al.
Published: (2025)
Causal Graph Learning via Distributional Invariance of Cause-Effect Relationship
by: Nguyen, Nang Hung, et al.
Published: (2026)
by: Nguyen, Nang Hung, et al.
Published: (2026)
Similar Items
-
Rethinking the Role of Temperature in Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026) -
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026) -
Enhancing Graph Neural Networks with Limited Labeled Data by Actively Distilling Knowledge from Large Language Models
by: Li, Quan, et al.
Published: (2024) -
Self-Supervised Quantization-Aware Knowledge Distillation
by: Zhao, Kaiqi, et al.
Published: (2024) -
RePCS: Diagnosing Data Memorization in LLM-Powered Retrieval-Augmented Generation
by: Anh, Le Vu, et al.
Published: (2025)