Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Luyang, Chen, Yongkai, Cai, Jiazhang, Ma, Ping, Zhong, Wenxuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning
by: Lu, Haoran, et al.
Published: (2026)
by: Lu, Haoran, et al.
Published: (2026)
Wahkon: A Statistically Principled Deep RKHS Superposition Network
by: Chen, Yongkai, et al.
Published: (2026)
by: Chen, Yongkai, et al.
Published: (2026)
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025)
by: Dalili, Seyed Arshan, et al.
Published: (2025)
M$^3$TN: Multi-gate Mixture-of-Experts based Multi-valued Treatment Network for Uplift Modeling
by: Sun, Zexu, et al.
Published: (2024)
by: Sun, Zexu, et al.
Published: (2024)
HOLOGRAPH: Active Causal Discovery via Sheaf-Theoretic Alignment of Large Language Model Priors
by: Kim, Hyunjun
Published: (2025)
by: Kim, Hyunjun
Published: (2025)
Interventional Causal Discovery in a Mixture of DAGs
by: Varıcı, Burak, et al.
Published: (2024)
by: Varıcı, Burak, et al.
Published: (2024)
Causal State Distillation for Explainable Reinforcement Learning
by: Lu, Wenhao, et al.
Published: (2023)
by: Lu, Wenhao, et al.
Published: (2023)
Mind the GAP: Improving Robustness to Subpopulation Shifts with Group-Aware Priors
by: Rudner, Tim G. J., et al.
Published: (2024)
by: Rudner, Tim G. J., et al.
Published: (2024)
Combining Priors with Experience: Confidence Calibration Based on Binomial Process Modeling
by: Dong, Jinzong, et al.
Published: (2024)
by: Dong, Jinzong, et al.
Published: (2024)
Provable Multi-Party Reinforcement Learning with Diverse Human Feedback
by: Zhong, Huiying, et al.
Published: (2024)
by: Zhong, Huiying, et al.
Published: (2024)
Efficient and Robust Knowledge Distillation from A Stronger Teacher Based on Correlation Matching
by: Niu, Wenqi, et al.
Published: (2024)
by: Niu, Wenqi, et al.
Published: (2024)
Joint Distribution-Informed Shapley Values for Sparse Counterfactual Explanations
by: You, Lei, et al.
Published: (2024)
by: You, Lei, et al.
Published: (2024)
Time Series Domain Adaptation via Latent Invariant Causal Mechanism
by: Cai, Ruichu, et al.
Published: (2025)
by: Cai, Ruichu, et al.
Published: (2025)
A Data-Driven Two-Phase Multi-Split Causal Ensemble Model for Time Series
by: Ma, Zhipeng, et al.
Published: (2024)
by: Ma, Zhipeng, et al.
Published: (2024)
Uncertainty Quantification for Prior-Data Fitted Networks using Martingale Posteriors
by: Nagler, Thomas, et al.
Published: (2025)
by: Nagler, Thomas, et al.
Published: (2025)
Exploring Multi-Modal Data with Tool-Augmented LLM Agents for Precise Causal Discovery
by: Shen, ChengAo, et al.
Published: (2024)
by: Shen, ChengAo, et al.
Published: (2024)
Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias
by: Li, Chao, et al.
Published: (2025)
by: Li, Chao, et al.
Published: (2025)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022)
by: Dong, Chengyu, et al.
Published: (2022)
Causal Diffusion Autoencoders: Toward Counterfactual Generation via Diffusion Probabilistic Models
by: Komanduri, Aneesh, et al.
Published: (2024)
by: Komanduri, Aneesh, et al.
Published: (2024)
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
by: Zhong, Yisheng, et al.
Published: (2026)
by: Zhong, Yisheng, et al.
Published: (2026)
Enhancing the Performance of Neural Networks Through Causal Discovery and Integration of Domain Knowledge
by: Zhang, Xiaoge, et al.
Published: (2023)
by: Zhang, Xiaoge, et al.
Published: (2023)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
by: Chen, Jinyin, et al.
Published: (2024)
by: Chen, Jinyin, et al.
Published: (2024)
Conformal Diffusion Models for Individual Treatment Effect Estimation and Inference
by: Cai, Hengrui, et al.
Published: (2024)
by: Cai, Hengrui, et al.
Published: (2024)
Generalized Independent Noise Condition for Estimating Causal Structure with Latent Variables
by: Xie, Feng, et al.
Published: (2023)
by: Xie, Feng, et al.
Published: (2023)
Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts
by: Fang, Luyang, et al.
Published: (2025)
by: Fang, Luyang, et al.
Published: (2025)
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
Can Large Language Models Help Experimental Design for Causal Discovery?
by: Li, Junyi, et al.
Published: (2025)
by: Li, Junyi, et al.
Published: (2025)
Localized Conformal Multi-Quantile Regression
by: Lu, Yuan
Published: (2024)
by: Lu, Yuan
Published: (2024)
The Role of Teacher Calibration in Knowledge Distillation
by: Kim, Suyoung, et al.
Published: (2025)
by: Kim, Suyoung, et al.
Published: (2025)
DCRMTA: Unbiased Causal Representation for Multi-touch Attribution
by: Tang, Jiaming
Published: (2024)
by: Tang, Jiaming
Published: (2024)
Multi-Domain Causal Discovery in Bijective Causal Models
by: Jalaldoust, Kasra, et al.
Published: (2025)
by: Jalaldoust, Kasra, et al.
Published: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
Regularized Multi-LLMs Collaboration for Enhanced Score-based Causal Discovery
by: Li, Xiaoxuan, et al.
Published: (2024)
by: Li, Xiaoxuan, et al.
Published: (2024)
Causal Layering via Conditional Entropy
by: Feigenbaum, Itai, et al.
Published: (2024)
by: Feigenbaum, Itai, et al.
Published: (2024)
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers
by: Nair, Lakshmi
Published: (2024)
by: Nair, Lakshmi
Published: (2024)
Weak Supervision Performance Evaluation via Partial Identification
by: Polo, Felipe Maia, et al.
Published: (2023)
by: Polo, Felipe Maia, et al.
Published: (2023)
Model-Free Assessment of Simulator Fidelity via Quantile Curves
by: Iyengar, Garud, et al.
Published: (2025)
by: Iyengar, Garud, et al.
Published: (2025)
Estimating and Mitigating the Congestion Effect of Curbside Pick-ups and Drop-offs: A Causal Inference Approach
by: Liu, Xiaohui, et al.
Published: (2022)
by: Liu, Xiaohui, et al.
Published: (2022)
FedMTFI: Feature Importance Based Optimized Multi Teacher Knowledge Distillation in Heterogeneous Federated Learning Environment
by: Shadin, Nazmus Shakib, et al.
Published: (2026)
by: Shadin, Nazmus Shakib, et al.
Published: (2026)
Similar Items
-
NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning
by: Lu, Haoran, et al.
Published: (2026) -
Wahkon: A Statistically Principled Deep RKHS Superposition Network
by: Chen, Yongkai, et al.
Published: (2026) -
Model Merging via Multi-Teacher Knowledge Distillation
by: Dalili, Seyed Arshan, et al.
Published: (2025) -
M$^3$TN: Multi-gate Mixture-of-Experts based Multi-valued Treatment Network for Uplift Modeling
by: Sun, Zexu, et al.
Published: (2024) -
HOLOGRAPH: Active Causal Discovery via Sheaf-Theoretic Alignment of Large Language Model Priors
by: Kim, Hyunjun
Published: (2025)