Boosting the Cross-Architecture Generalization of Dataset Distillation through an Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Lirui, Zhang, Yuxin, Chao, Fei, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Efficient Automatic Self-Pruning of Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Improve Cross-Architecture Generalization on Dataset Distillation
by: Zhou, Binglin, et al.
Published: (2024)
by: Zhou, Binglin, et al.
Published: (2024)
Towards Mitigating Architecture Overfitting on Distilled Datasets
by: Zhong, Xuyang, et al.
Published: (2023)
by: Zhong, Xuyang, et al.
Published: (2023)
Learning under Singularity: An Information Criterion improving WBIC and sBIC
by: Liu, Lirui, et al.
Published: (2024)
by: Liu, Lirui, et al.
Published: (2024)
Tailoring Instructions to Student's Learning Levels Boosts Knowledge Distillation
by: Ren, Yuxin, et al.
Published: (2023)
by: Ren, Yuxin, et al.
Published: (2023)
Towards Trustworthy Dataset Distillation
by: Ma, Shijie, et al.
Published: (2023)
by: Ma, Shijie, et al.
Published: (2023)
De-Anonymization at Scale via Tournament-Style Attribution
by: Zhang, Lirui, et al.
Published: (2026)
by: Zhang, Lirui, et al.
Published: (2026)
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Information-Theoretic Criteria for Knowledge Distillation in Multimodal Learning
by: Xie, Rongrong, et al.
Published: (2025)
by: Xie, Rongrong, et al.
Published: (2025)
Distilling Long-tailed Datasets
by: Zhao, Zhenghao, et al.
Published: (2024)
by: Zhao, Zhenghao, et al.
Published: (2024)
TimeDistill: Efficient Long-Term Time Series Forecasting with MLP via Cross-Architecture Distillation
by: Ni, Juntong, et al.
Published: (2025)
by: Ni, Juntong, et al.
Published: (2025)
Polybasic Speculative Decoding Through a Theoretical Perspective
by: Wang, Ruilin, et al.
Published: (2025)
by: Wang, Ruilin, et al.
Published: (2025)
Cross-Dataset Generalization in Deep Learning
by: Zhang, Xuyu, et al.
Published: (2024)
by: Zhang, Xuyu, et al.
Published: (2024)
Prioritize Alignment in Dataset Distillation
by: Li, Zekai, et al.
Published: (2024)
by: Li, Zekai, et al.
Published: (2024)
What Can RL Bring to VLA Generalization? An Empirical Study
by: Liu, Jijia, et al.
Published: (2025)
by: Liu, Jijia, et al.
Published: (2025)
Younger: The First Dataset for Artificial Intelligence-Generated Neural Network Architecture
by: Yang, Zhengxin, et al.
Published: (2024)
by: Yang, Zhengxin, et al.
Published: (2024)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
by: Singh, Aasheesh, et al.
Published: (2025)
by: Singh, Aasheesh, et al.
Published: (2025)
DNAD: Differentiable Neural Architecture Distillation
by: Rao, Xuan, et al.
Published: (2025)
by: Rao, Xuan, et al.
Published: (2025)
Dynamic Low-Rank Sparse Adaptation for Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Fair Dataset Distillation via Cross-Group Barycenter Alignment
by: Moslemi, Mohammad Hossein, et al.
Published: (2026)
by: Moslemi, Mohammad Hossein, et al.
Published: (2026)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Generalization Gaps in Political Fake News Detection: An Empirical Study on the LIAR Dataset
by: Hasan, S Mahmudul, et al.
Published: (2025)
by: Hasan, S Mahmudul, et al.
Published: (2025)
A Systematic Empirical Study of Grokking: Depth, Architecture, Activation, and Regularization
by: Manir, Shalima Binta, et al.
Published: (2026)
by: Manir, Shalima Binta, et al.
Published: (2026)
Generalized Kernel Inducing Points by Duality Gap for Dataset Distillation
by: Aoyama, Tatsuya, et al.
Published: (2025)
by: Aoyama, Tatsuya, et al.
Published: (2025)
PRISM: Diversifying Dataset Distillation by Decoupling Architectural Priors
by: Moser, Brian B., et al.
Published: (2025)
by: Moser, Brian B., et al.
Published: (2025)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
MMICT: Boosting Multi-Modal Fine-Tuning with In-Context Examples
by: Chen, Tao, et al.
Published: (2023)
by: Chen, Tao, et al.
Published: (2023)
AdaGMLP: AdaBoosting GNN-to-MLP Knowledge Distillation
by: Lu, Weigang, et al.
Published: (2024)
by: Lu, Weigang, et al.
Published: (2024)
What is Dataset Distillation Learning?
by: Yang, William, et al.
Published: (2024)
by: Yang, William, et al.
Published: (2024)
Structure-Attribute Transformations with Markov Chain Boost Graph Domain Adaptation
by: Liu, Zhen, et al.
Published: (2025)
by: Liu, Zhen, et al.
Published: (2025)
Revisiting Agnostic Boosting
by: da Cunha, Arthur, et al.
Published: (2025)
by: da Cunha, Arthur, et al.
Published: (2025)
IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning
by: Qin, Yihao, et al.
Published: (2026)
by: Qin, Yihao, et al.
Published: (2026)
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
by: Zhang, Xingxuan, et al.
Published: (2025)
by: Zhang, Xingxuan, et al.
Published: (2025)
An Empirical Investigation into the Effect of Parameter Choices in Knowledge Distillation
by: Sultan, Md Arafat, et al.
Published: (2024)
by: Sultan, Md Arafat, et al.
Published: (2024)
FedDTG:Federated Data-Free Knowledge Distillation via Three-Player Generative Adversarial Networks
by: Gao, Lingzhi, et al.
Published: (2022)
by: Gao, Lingzhi, et al.
Published: (2022)
Training Diverse Graph Experts for Ensembles: A Systematic Empirical Study
by: Deng, Gangda, et al.
Published: (2025)
by: Deng, Gangda, et al.
Published: (2025)
Generative Dataset Distillation Based on Self-knowledge Distillation
by: Li, Longzhen, et al.
Published: (2025)
by: Li, Longzhen, et al.
Published: (2025)
SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement Learning
by: Luo, Lirui, et al.
Published: (2026)
by: Luo, Lirui, et al.
Published: (2026)
Diffusion Models as Dataset Distillation Priors
by: Su, Duo, et al.
Published: (2025)
by: Su, Duo, et al.
Published: (2025)
Similar Items
-
Towards Efficient Automatic Self-Pruning of Large Language Models
by: Huang, Weizhong, et al.
Published: (2025) -
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
by: Huang, Weizhong, et al.
Published: (2025) -
Improve Cross-Architecture Generalization on Dataset Distillation
by: Zhou, Binglin, et al.
Published: (2024) -
Towards Mitigating Architecture Overfitting on Distilled Datasets
by: Zhong, Xuyang, et al.
Published: (2023) -
Learning under Singularity: An Information Criterion improving WBIC and sBIC
by: Liu, Lirui, et al.
Published: (2024)