Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Rang, Miao, Bi, Zhenni, Zhou, Hang, Chen, Hanting, Xiao, An, Guo, Tianyu, Han, Kai, Chen, Xinghao, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
by: Rang, Miao, et al.
Published: (2026)
by: Rang, Miao, et al.
Published: (2026)
An Empirical Study of Scaling Law for OCR
by: Rang, Miao, et al.
Published: (2023)
by: Rang, Miao, et al.
Published: (2023)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
An Empirical Study of World Model Quantization
by: Fu, Zhongqian, et al.
Published: (2026)
by: Fu, Zhongqian, et al.
Published: (2026)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
by: Guo, Jianyuan, et al.
Published: (2024)
by: Guo, Jianyuan, et al.
Published: (2024)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
by: Guo, Jialong, et al.
Published: (2024)
by: Guo, Jialong, et al.
Published: (2024)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
by: Nie, Ying, et al.
Published: (2025)
by: Nie, Ying, et al.
Published: (2025)
Nexus: Higher-Order Attention Mechanisms in Transformers
by: Chen, Hanting, et al.
Published: (2025)
by: Chen, Hanting, et al.
Published: (2025)
DECO: Unleashing the Potential of ConvNets for Query-based Detection and Segmentation
by: Chen, Xinghao, et al.
Published: (2023)
by: Chen, Xinghao, et al.
Published: (2023)
Distilling Semantic Priors from SAM to Efficient Image Restoration Models
by: Zhang, Quan, et al.
Published: (2024)
by: Zhang, Quan, et al.
Published: (2024)
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
by: Bi, Zhenni, et al.
Published: (2024)
by: Bi, Zhenni, et al.
Published: (2024)
SAM-TTT: Segment Anything Model via Reverse Parameter Configuration and Test-Time Training for Camouflaged Object Detection
by: Yu, Zhenni, et al.
Published: (2025)
by: Yu, Zhenni, et al.
Published: (2025)
Knowledge Distillation with Multi-granularity Mixture of Priors for Image Super-Resolution
by: Li, Simiao, et al.
Published: (2024)
by: Li, Simiao, et al.
Published: (2024)
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
by: Zhai, Yingjie, et al.
Published: (2024)
by: Zhai, Yingjie, et al.
Published: (2024)
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
by: Ni, Zhenliang, et al.
Published: (2024)
by: Ni, Zhenliang, et al.
Published: (2024)
From Sequential to Spatial: Reordering Autoregression for Efficient Visual Generation
by: Wang, Siyang, et al.
Published: (2025)
by: Wang, Siyang, et al.
Published: (2025)
Data Upcycling Knowledge Distillation for Image Super-Resolution
by: Zhang, Yun, et al.
Published: (2023)
by: Zhang, Yun, et al.
Published: (2023)
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
by: Ren, Weining, et al.
Published: (2025)
by: Ren, Weining, et al.
Published: (2025)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
by: Shu, Han, et al.
Published: (2023)
by: Shu, Han, et al.
Published: (2023)
GhostNetV3: Exploring the Training Strategies for Compact Models
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
ParameterNet: Parameters Are All You Need
by: Han, Kai, et al.
Published: (2023)
by: Han, Kai, et al.
Published: (2023)
Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
by: Sun, Jingchen, et al.
Published: (2026)
by: Sun, Jingchen, et al.
Published: (2026)
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
by: Huang, Mouxiao, et al.
Published: (2025)
by: Huang, Mouxiao, et al.
Published: (2025)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
by: Han, Kai, et al.
Published: (2024)
by: Han, Kai, et al.
Published: (2024)
IPT-V2: Efficient Image Processing Transformer using Hierarchical Attentions
by: Tu, Zhijun, et al.
Published: (2024)
by: Tu, Zhijun, et al.
Published: (2024)
DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution
by: Cai, Miaomiao, et al.
Published: (2025)
by: Cai, Miaomiao, et al.
Published: (2025)
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping
by: Ding, Ning, et al.
Published: (2025)
by: Ding, Ning, et al.
Published: (2025)
SlimLLM: Accurate Structured Pruning for Large Language Models
by: Guo, Jialong, et al.
Published: (2025)
by: Guo, Jialong, et al.
Published: (2025)
Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge Retention
by: Lin, Wenye, et al.
Published: (2026)
by: Lin, Wenye, et al.
Published: (2026)
Weakly Supervised Change Detection via Knowledge Distillation and Multiscale Sigmoid Inference
by: Lu, Binghao, et al.
Published: (2024)
by: Lu, Binghao, et al.
Published: (2024)
Lightweight Model Pre-training via Language Guided Knowledge Distillation
by: Li, Mingsheng, et al.
Published: (2024)
by: Li, Mingsheng, et al.
Published: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer
by: Jia, Ding, et al.
Published: (2024)
by: Jia, Ding, et al.
Published: (2024)
Efficient Diffusion Training via Min-SNR Weighting Strategy
by: Hang, Tiankai, et al.
Published: (2023)
by: Hang, Tiankai, et al.
Published: (2023)
Similar Items
-
Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
by: Rang, Miao, et al.
Published: (2026) -
An Empirical Study of Scaling Law for OCR
by: Rang, Miao, et al.
Published: (2023) -
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
by: Rang, Miao, et al.
Published: (2025) -
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
by: He, Wei, et al.
Published: (2025) -
An Empirical Study of World Model Quantization
by: Fu, Zhongqian, et al.
Published: (2026)