Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Chengming, Wu, Haolun, Li, Xuan, Ma, Chen, Chen, Xi, Yan, Jun, Wang, Boyu, Liu, Xue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
por: Li, Tong, et al.
Publicado: (2025)
por: Li, Tong, et al.
Publicado: (2025)
When Less is More: The LLM Scaling Paradox in Context Compression
por: Guo, Ruishan, et al.
Publicado: (2026)
por: Guo, Ruishan, et al.
Publicado: (2026)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
por: Fang, Luyang, et al.
Publicado: (2026)
por: Fang, Luyang, et al.
Publicado: (2026)
Knowledge Distillation with Adapted Weight
por: Wu, Sirong, et al.
Publicado: (2025)
por: Wu, Sirong, et al.
Publicado: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
por: Shen, Yuhao, et al.
Publicado: (2026)
por: Shen, Yuhao, et al.
Publicado: (2026)
Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation
por: Zheng, Yilun, et al.
Publicado: (2025)
por: Zheng, Yilun, et al.
Publicado: (2025)
Quantize What Counts: More for Keys, Less for Values
por: Hariri, Mohsen, et al.
Publicado: (2025)
por: Hariri, Mohsen, et al.
Publicado: (2025)
An Overview of Machine Learning-Enabled Optimization for Reconfigurable Intelligent Surfaces-Aided 6G Networks: From Reinforcement Learning to Large Language Models
por: Zhou, Hao, et al.
Publicado: (2024)
por: Zhou, Hao, et al.
Publicado: (2024)
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
por: Wu, Lirong, et al.
Publicado: (2024)
por: Wu, Lirong, et al.
Publicado: (2024)
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
por: Li, Hui, et al.
Publicado: (2025)
por: Li, Hui, et al.
Publicado: (2025)
Cut Less, Fold More: Model Compression through the Lens of Projection Geometry
por: Saukh, Olga, et al.
Publicado: (2026)
por: Saukh, Olga, et al.
Publicado: (2026)
How to Backdoor the Knowledge Distillation
por: Wu, Chen, et al.
Publicado: (2025)
por: Wu, Chen, et al.
Publicado: (2025)
Less is More: Towards Simple Graph Contrastive Learning
por: Zhao, Yanan, et al.
Publicado: (2025)
por: Zhao, Yanan, et al.
Publicado: (2025)
Transformer Multivariate Forecasting: Less is More?
por: Xu, Jingjing, et al.
Publicado: (2023)
por: Xu, Jingjing, et al.
Publicado: (2023)
Moirai 2.0: When Less Is More for Time Series Forecasting
por: Liu, Chenghao, et al.
Publicado: (2025)
por: Liu, Chenghao, et al.
Publicado: (2025)
ViRN: Variational Inference and Distribution Trilateration for Long-Tailed Continual Representation Learning
por: Dai, Hao, et al.
Publicado: (2025)
por: Dai, Hao, et al.
Publicado: (2025)
Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias
por: Li, Chao, et al.
Publicado: (2025)
por: Li, Chao, et al.
Publicado: (2025)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
por: Panigrahi, Abhishek, et al.
Publicado: (2025)
por: Panigrahi, Abhishek, et al.
Publicado: (2025)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
por: Wu, Shaojin, et al.
Publicado: (2025)
por: Wu, Shaojin, et al.
Publicado: (2025)
Making Recommender Systems More Knowledgeable: A Framework to Incorporate Side Information
por: Jiang, Yukun, et al.
Publicado: (2024)
por: Jiang, Yukun, et al.
Publicado: (2024)
How to Train the Teacher Model for Effective Knowledge Distillation
por: Hamidi, Shayan Mohajer, et al.
Publicado: (2024)
por: Hamidi, Shayan Mohajer, et al.
Publicado: (2024)
Learning More with Less: A Generalizable, Self-Supervised Framework for Privacy-Preserving Capacity Estimation with EV Charging Data
por: Arunan, Anushiya, et al.
Publicado: (2025)
por: Arunan, Anushiya, et al.
Publicado: (2025)
Less is More: Efficient Weight Farcasting with 1-Layer Neural Network
por: Shou, Xiao, et al.
Publicado: (2025)
por: Shou, Xiao, et al.
Publicado: (2025)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
por: Dong, Chengyu, et al.
Publicado: (2022)
por: Dong, Chengyu, et al.
Publicado: (2022)
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
por: Binici, Kuluhan, et al.
Publicado: (2024)
por: Binici, Kuluhan, et al.
Publicado: (2024)
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
por: Liu, Jinzhe, et al.
Publicado: (2025)
por: Liu, Jinzhe, et al.
Publicado: (2025)
Less Approximates More: Harmonizing Performance and Confidence Faithfulness via Hybrid Post-Training for High-Stakes Tasks
por: Ma, Haokai, et al.
Publicado: (2026)
por: Ma, Haokai, et al.
Publicado: (2026)
Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning
por: Zhao, Lifan, et al.
Publicado: (2025)
por: Zhao, Lifan, et al.
Publicado: (2025)
Online Knowledge Distillation with Reward Guidance
por: Jia, Chen
Publicado: (2025)
por: Jia, Chen
Publicado: (2025)
Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities
por: Zhou, Hao, et al.
Publicado: (2024)
por: Zhou, Hao, et al.
Publicado: (2024)
SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines
por: Morad, Itai, et al.
Publicado: (2026)
por: Morad, Itai, et al.
Publicado: (2026)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
por: Chen, Jinyin, et al.
Publicado: (2024)
por: Chen, Jinyin, et al.
Publicado: (2024)
LIMR: Less is More for RL Scaling
por: Li, Xuefeng, et al.
Publicado: (2025)
por: Li, Xuefeng, et al.
Publicado: (2025)
Diffusion Models as Dataset Distillation Priors
por: Su, Duo, et al.
Publicado: (2025)
por: Su, Duo, et al.
Publicado: (2025)
Model Merging via Multi-Teacher Knowledge Distillation
por: Dalili, Seyed Arshan, et al.
Publicado: (2025)
por: Dalili, Seyed Arshan, et al.
Publicado: (2025)
Less is More: Adaptive Coverage for Synthetic Training Data
por: Tavakkol, Sasan, et al.
Publicado: (2025)
por: Tavakkol, Sasan, et al.
Publicado: (2025)
Self-Ablating Transformers: More Interpretability, Less Sparsity
por: Ferrao, Jeremias, et al.
Publicado: (2025)
por: Ferrao, Jeremias, et al.
Publicado: (2025)
Stragglers Can Contribute More: Uncertainty-Aware Distillation for Asynchronous Federated Learning
por: Wang, Yujia, et al.
Publicado: (2025)
por: Wang, Yujia, et al.
Publicado: (2025)
Knowledge Distillation Based on Transformed Teacher Matching
por: Zheng, Kaixiang, et al.
Publicado: (2024)
por: Zheng, Kaixiang, et al.
Publicado: (2024)
Continual Policy Distillation from Distributed Reinforcement Learning Teachers
por: Li, Yuxuan, et al.
Publicado: (2026)
por: Li, Yuxuan, et al.
Publicado: (2026)
Ejemplares similares
-
Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
por: Li, Tong, et al.
Publicado: (2025) -
When Less is More: The LLM Scaling Paradox in Context Compression
por: Guo, Ruishan, et al.
Publicado: (2026) -
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
por: Fang, Luyang, et al.
Publicado: (2026) -
Knowledge Distillation with Adapted Weight
por: Wu, Sirong, et al.
Publicado: (2025) -
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
por: Shen, Yuhao, et al.
Publicado: (2026)