From Teacher to Student: Tracking Memorization Through Model Distillation
Fuente:
arXiv
Guardado en:
| Autor principal: | Singh, Simardeep |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Memorization of Consistency Distillation for Diffusion Models
por: Jiang, Bingqing, et al.
Publicado: (2026)
por: Jiang, Bingqing, et al.
Publicado: (2026)
Membership and Memorization in LLM Knowledge Distillation
por: Zhang, Ziqi, et al.
Publicado: (2025)
por: Zhang, Ziqi, et al.
Publicado: (2025)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
por: Dong, Chengyu, et al.
Publicado: (2022)
por: Dong, Chengyu, et al.
Publicado: (2022)
Teaching the Teacher: The Role of Teacher-Student Smoothness Alignment in Genetic Programming-based Symbolic Distillation
por: Dhar, Soumyadeep, et al.
Publicado: (2025)
por: Dhar, Soumyadeep, et al.
Publicado: (2025)
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
por: Binici, Kuluhan, et al.
Publicado: (2024)
por: Binici, Kuluhan, et al.
Publicado: (2024)
Model Merging via Multi-Teacher Knowledge Distillation
por: Dalili, Seyed Arshan, et al.
Publicado: (2025)
por: Dalili, Seyed Arshan, et al.
Publicado: (2025)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
por: Lee, Byung-Kwan, et al.
Publicado: (2025)
por: Lee, Byung-Kwan, et al.
Publicado: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
por: Bayat, Reza, et al.
Publicado: (2024)
por: Bayat, Reza, et al.
Publicado: (2024)
Memorization Sinks: Isolating Memorization during LLM Training
por: Ghosal, Gaurav R., et al.
Publicado: (2025)
por: Ghosal, Gaurav R., et al.
Publicado: (2025)
On Teacher Hacking in Language Model Distillation
por: Tiapkin, Daniil, et al.
Publicado: (2025)
por: Tiapkin, Daniil, et al.
Publicado: (2025)
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher Model
por: Chen, Jinyin, et al.
Publicado: (2024)
por: Chen, Jinyin, et al.
Publicado: (2024)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
por: Xu, Yuanda, et al.
Publicado: (2026)
por: Xu, Yuanda, et al.
Publicado: (2026)
Linear Projections of Teacher Embeddings for Few-Class Distillation
por: Loo, Noel, et al.
Publicado: (2024)
por: Loo, Noel, et al.
Publicado: (2024)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
por: Behrouz, Ali, et al.
Publicado: (2025)
por: Behrouz, Ali, et al.
Publicado: (2025)
Analyzing Memorization in Large Language Models through the Lens of Model Attribution
por: Menta, Tarun Ram, et al.
Publicado: (2025)
por: Menta, Tarun Ram, et al.
Publicado: (2025)
Mitigating Memorization In Language Models
por: Sakarvadia, Mansi, et al.
Publicado: (2024)
por: Sakarvadia, Mansi, et al.
Publicado: (2024)
Generalizability of Memorization Neural Networks
por: Yu, Lijia, et al.
Publicado: (2024)
por: Yu, Lijia, et al.
Publicado: (2024)
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
por: Fang, Luyang, et al.
Publicado: (2026)
por: Fang, Luyang, et al.
Publicado: (2026)
Beyond Frequency: The Role of Redundancy in Large Language Model Memorization
por: Zhang, Jie, et al.
Publicado: (2025)
por: Zhang, Jie, et al.
Publicado: (2025)
Memorization Control in Diffusion Models from Denoising-centric Perspective
por: Vu, Thuy Phuong, et al.
Publicado: (2026)
por: Vu, Thuy Phuong, et al.
Publicado: (2026)
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
por: Singh, Simardeep, et al.
Publicado: (2026)
por: Singh, Simardeep, et al.
Publicado: (2026)
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
por: Zhong, Yisheng, et al.
Publicado: (2026)
por: Zhong, Yisheng, et al.
Publicado: (2026)
Memorization in deep learning: A survey
por: Wei, Jiaheng, et al.
Publicado: (2024)
por: Wei, Jiaheng, et al.
Publicado: (2024)
Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models
por: Hintersdorf, Dominik, et al.
Publicado: (2024)
por: Hintersdorf, Dominik, et al.
Publicado: (2024)
Understanding and Mitigating Memorization in Generative Models via Sharpness of Probability Landscapes
por: Jeon, Dongjae, et al.
Publicado: (2024)
por: Jeon, Dongjae, et al.
Publicado: (2024)
Detecting Memorization in Large Language Models
por: Slonski, Eduardo
Publicado: (2024)
por: Slonski, Eduardo
Publicado: (2024)
Teacher-Student Learning on Complexity in Intelligent Routing
por: Pi, Shu-Ting, et al.
Publicado: (2024)
por: Pi, Shu-Ting, et al.
Publicado: (2024)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
por: Qin, Zongyue, et al.
Publicado: (2025)
por: Qin, Zongyue, et al.
Publicado: (2025)
Quantifying Memorization and Parametric Response Rates in Retrieval-Augmented Vision-Language Models
por: Carragher, Peter, et al.
Publicado: (2025)
por: Carragher, Peter, et al.
Publicado: (2025)
How Do Flow Matching Models Memorize and Generalize in Sample Data Subspaces?
por: Gao, Weiguo, et al.
Publicado: (2024)
por: Gao, Weiguo, et al.
Publicado: (2024)
Re-understanding Graph Unlearning through Memorization
por: Ding, Pengfei, et al.
Publicado: (2026)
por: Ding, Pengfei, et al.
Publicado: (2026)
Batch Normalization Amplifies Memorization and Privacy Risks
por: Doan, Ngoc Phu, et al.
Publicado: (2026)
por: Doan, Ngoc Phu, et al.
Publicado: (2026)
On Memorization in Diffusion Models
por: Gu, Xiangming, et al.
Publicado: (2023)
por: Gu, Xiangming, et al.
Publicado: (2023)
Group Relative Knowledge Distillation: Learning from Teacher's Relational Inductive Bias
por: Li, Chao, et al.
Publicado: (2025)
por: Li, Chao, et al.
Publicado: (2025)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
por: Liu, Xiaogeng, et al.
Publicado: (2026)
por: Liu, Xiaogeng, et al.
Publicado: (2026)
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers
por: Nair, Lakshmi
Publicado: (2024)
por: Nair, Lakshmi
Publicado: (2024)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
por: Singh, Karan, et al.
Publicado: (2026)
por: Singh, Karan, et al.
Publicado: (2026)
Exploring Memorization in Fine-tuned Language Models
por: Zeng, Shenglai, et al.
Publicado: (2023)
por: Zeng, Shenglai, et al.
Publicado: (2023)
Are Large Language Models Memorizing Bug Benchmarks?
por: Ramos, Daniel, et al.
Publicado: (2024)
por: Ramos, Daniel, et al.
Publicado: (2024)
DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation
por: Wang, Zexin, et al.
Publicado: (2025)
por: Wang, Zexin, et al.
Publicado: (2025)
Ejemplares similares
-
On the Memorization of Consistency Distillation for Diffusion Models
por: Jiang, Bingqing, et al.
Publicado: (2026) -
Membership and Memorization in LLM Knowledge Distillation
por: Zhang, Ziqi, et al.
Publicado: (2025) -
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
por: Dong, Chengyu, et al.
Publicado: (2022) -
Teaching the Teacher: The Role of Teacher-Student Smoothness Alignment in Genetic Programming-based Symbolic Distillation
por: Dhar, Soumyadeep, et al.
Publicado: (2025) -
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures
por: Binici, Kuluhan, et al.
Publicado: (2024)