Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-Training of Deep Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Joshi, Siddharth, Ni, Jiayi, Mirzasoleiman, Baharan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
di: Joshi, Siddharth, et al.
Pubblicazione: (2023)
di: Joshi, Siddharth, et al.
Pubblicazione: (2023)
Investigating the Benefits of Projection Head for Representation Learning
di: Xue, Yihao, et al.
Pubblicazione: (2024)
di: Xue, Yihao, et al.
Pubblicazione: (2024)
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
di: Joshi, Siddharth, et al.
Pubblicazione: (2024)
di: Joshi, Siddharth, et al.
Pubblicazione: (2024)
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
di: Xue, Yihao, et al.
Pubblicazione: (2023)
di: Xue, Yihao, et al.
Pubblicazione: (2023)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
di: Javanmard, Adel, et al.
Pubblicazione: (2026)
di: Javanmard, Adel, et al.
Pubblicazione: (2026)
Graph Contrastive Learning under Heterophily via Graph Filters
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
Challenges and Opportunities in Improving Worst-Group Generalization in Presence of Spurious Features
di: Joshi, Siddharth, et al.
Pubblicazione: (2023)
di: Joshi, Siddharth, et al.
Pubblicazione: (2023)
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
di: Huang, Jianhao, et al.
Pubblicazione: (2026)
di: Huang, Jianhao, et al.
Pubblicazione: (2026)
Understanding the Role of Training Data in Test-Time Scaling
di: Javanmard, Adel, et al.
Pubblicazione: (2025)
di: Javanmard, Adel, et al.
Pubblicazione: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
di: Cazenavette, George, et al.
Pubblicazione: (2025)
di: Cazenavette, George, et al.
Pubblicazione: (2025)
Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision
di: Xue, Yihao, et al.
Pubblicazione: (2026)
di: Xue, Yihao, et al.
Pubblicazione: (2026)
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
di: Yang, Wenhan, et al.
Pubblicazione: (2023)
Investigating the Impact of Model Width and Density on Generalization in Presence of Label Noise
di: Xue, Yihao, et al.
Pubblicazione: (2022)
di: Xue, Yihao, et al.
Pubblicazione: (2022)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
di: Xue, Yihao, et al.
Pubblicazione: (2025)
di: Xue, Yihao, et al.
Pubblicazione: (2025)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
di: Joshi, Siddharth, et al.
Pubblicazione: (2025)
di: Joshi, Siddharth, et al.
Pubblicazione: (2025)
Self-Supervised Dataset Distillation for Transfer Learning
di: Lee, Dong Bok, et al.
Pubblicazione: (2023)
di: Lee, Dong Bok, et al.
Pubblicazione: (2023)
Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution Generalization
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
di: Yang, Yu, et al.
Pubblicazione: (2023)
di: Yang, Yu, et al.
Pubblicazione: (2023)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
di: Gangavarapu, Tushaar, et al.
Pubblicazione: (2026)
di: Gangavarapu, Tushaar, et al.
Pubblicazione: (2026)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
di: Yang, Yu, et al.
Pubblicazione: (2024)
di: Yang, Yu, et al.
Pubblicazione: (2024)
Practical Insights into Knowledge Distillation for Pre-Trained Models
di: Alballa, Norah, et al.
Pubblicazione: (2024)
di: Alballa, Norah, et al.
Pubblicazione: (2024)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
di: Dong, Chengyu, et al.
Pubblicazione: (2022)
di: Dong, Chengyu, et al.
Pubblicazione: (2022)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
Efficient Knowledge Injection in LLMs via Self-Distillation
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
Self-Supervised Quantization-Aware Knowledge Distillation
di: Zhao, Kaiqi, et al.
Pubblicazione: (2024)
di: Zhao, Kaiqi, et al.
Pubblicazione: (2024)
PPG-Distill: Efficient Photoplethysmography Signals Analysis via Foundation Model Distillation
di: Ni, Juntong, et al.
Pubblicazione: (2025)
di: Ni, Juntong, et al.
Pubblicazione: (2025)
Distill to Delete: Unlearning in Graph Networks with Knowledge Distillation
di: Sinha, Yash, et al.
Pubblicazione: (2023)
di: Sinha, Yash, et al.
Pubblicazione: (2023)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
di: Xue, Yihao, et al.
Pubblicazione: (2023)
di: Xue, Yihao, et al.
Pubblicazione: (2023)
Convex Distillation: Efficient Compression of Deep Networks via Convex Optimization
di: Varshney, Prateek, et al.
Pubblicazione: (2024)
di: Varshney, Prateek, et al.
Pubblicazione: (2024)
Knowledge Distillation with Training Wheels
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training
di: Tang, Yuting, et al.
Pubblicazione: (2025)
di: Tang, Yuting, et al.
Pubblicazione: (2025)
How Transformers Learn to Plan via Multi-Token Prediction
di: Huang, Jianhao, et al.
Pubblicazione: (2026)
di: Huang, Jianhao, et al.
Pubblicazione: (2026)
Algorithmic Guarantees for Distilling Supervised and Offline RL Datasets
di: Gupta, Aaryan, et al.
Pubblicazione: (2025)
di: Gupta, Aaryan, et al.
Pubblicazione: (2025)
TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation
di: Ni, Juntong, et al.
Pubblicazione: (2025)
di: Ni, Juntong, et al.
Pubblicazione: (2025)
The Role of Masking for Efficient Supervised Knowledge Distillation of Vision Transformers
di: Son, Seungwoo, et al.
Pubblicazione: (2023)
di: Son, Seungwoo, et al.
Pubblicazione: (2023)
HERO: Heterogeneous Continual Graph Learning via Meta-Knowledge Distillation
di: Sun, Guiquan, et al.
Pubblicazione: (2025)
di: Sun, Guiquan, et al.
Pubblicazione: (2025)
A Survey on Pre-Trained Diffusion Model Distillations
di: Fan, Xuhui, et al.
Pubblicazione: (2025)
di: Fan, Xuhui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
di: Joshi, Siddharth, et al.
Pubblicazione: (2023) -
Investigating the Benefits of Projection Head for Representation Learning
di: Xue, Yihao, et al.
Pubblicazione: (2024) -
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
di: Joshi, Siddharth, et al.
Pubblicazione: (2024) -
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
di: Xue, Yihao, et al.
Pubblicazione: (2023) -
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
di: Javanmard, Adel, et al.
Pubblicazione: (2026)