Asymmetric Masked Distillation for Pre-Training Small Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhiyu, Huang, Bingkun, Xing, Sen, Wu, Gangshan, Qiao, Yu, Wang, Limin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
by: Zhang, Jiaming, et al.
Published: (2023)
by: Zhang, Jiaming, et al.
Published: (2023)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023)
by: Cui, Yutao, et al.
Published: (2023)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
by: Zhu, Chenhui, et al.
Published: (2025)
by: Zhu, Chenhui, et al.
Published: (2025)
Data-efficient Event Camera Pre-training via Disentangled Masked Modeling
by: Huang, Zhenpeng, et al.
Published: (2024)
by: Huang, Zhenpeng, et al.
Published: (2024)
STMixer: A One-Stage Sparse Action Detector
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Open-Vocabulary Spatio-Temporal Action Detection
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Dual DETRs for Multi-Label Temporal Action Detection
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
Pure-Pass: Fine-Grained, Adaptive Masking for Dynamic Token-Mixing Routing in Lightweight Image Super-Resolution
by: Wu, Junyu, et al.
Published: (2025)
by: Wu, Junyu, et al.
Published: (2025)
Efficient Test-Time Prompt Tuning for Vision-Language Models
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
by: Luo, Xingyu, et al.
Published: (2026)
by: Luo, Xingyu, et al.
Published: (2026)
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
by: Wu, Biao, et al.
Published: (2024)
by: Wu, Biao, et al.
Published: (2024)
Harvest Video Foundation Models via Efficient Post-Pretraining
by: Li, Yizhuo, et al.
Published: (2023)
by: Li, Yizhuo, et al.
Published: (2023)
MiM: Mask in Mask Self-Supervised Pre-Training for 3D Medical Image Analysis
by: Zhuang, Jiaxin, et al.
Published: (2024)
by: Zhuang, Jiaxin, et al.
Published: (2024)
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
by: Zheng, Xuhui, et al.
Published: (2025)
by: Zheng, Xuhui, et al.
Published: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
by: Pei, Baoqi, et al.
Published: (2024)
by: Pei, Baoqi, et al.
Published: (2024)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
by: Zhu, Haoyi, et al.
Published: (2023)
by: Zhu, Haoyi, et al.
Published: (2023)
Pre-Trained Masked Image Model for Mobile Robot Navigation
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
by: Li, Wenyu, et al.
Published: (2025)
by: Li, Wenyu, et al.
Published: (2025)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
by: Röhrich, Nikolai, et al.
Published: (2025)
by: Röhrich, Nikolai, et al.
Published: (2025)
Multi-Modal Soccer Scene Analysis with Masked Pre-Training
by: Peral, Marc, et al.
Published: (2025)
by: Peral, Marc, et al.
Published: (2025)
SOEDiff: Efficient Distillation for Small Object Editing
by: Wu, Yiming, et al.
Published: (2024)
by: Wu, Yiming, et al.
Published: (2024)
An Experimental Study on Exploring Strong Lightweight Vision Transformers via Masked Image Modeling Pre-Training
by: Gao, Jin, et al.
Published: (2024)
by: Gao, Jin, et al.
Published: (2024)
CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
by: Wu, Xiaoxue, et al.
Published: (2025)
by: Wu, Xiaoxue, et al.
Published: (2025)
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
by: Fan, Jiawei, et al.
Published: (2026)
by: Fan, Jiawei, et al.
Published: (2026)
ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal Prediction
by: Zhang, Wenxuan, et al.
Published: (2023)
by: Zhang, Wenxuan, et al.
Published: (2023)
Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching
by: Xia, Qianxin, et al.
Published: (2026)
by: Xia, Qianxin, et al.
Published: (2026)
ComKD-CLIP: Comprehensive Knowledge Distillation for Contrastive Language-Image Pre-traning Model
by: Chen, Yifan, et al.
Published: (2024)
by: Chen, Yifan, et al.
Published: (2024)
Reprogramming Distillation for Medical Foundation Models
by: Zhou, Yuhang, et al.
Published: (2024)
by: Zhou, Yuhang, et al.
Published: (2024)
Lightweight Model Pre-training via Language Guided Knowledge Distillation
by: Li, Mingsheng, et al.
Published: (2024)
by: Li, Mingsheng, et al.
Published: (2024)
Deep Reprogramming Distillation for Medical Foundation Models
by: Du, Siyuan, et al.
Published: (2026)
by: Du, Siyuan, et al.
Published: (2026)
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
by: Chen, Yitong, et al.
Published: (2025)
by: Chen, Yitong, et al.
Published: (2025)
Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation
by: Yang, Yuchen, et al.
Published: (2023)
by: Yang, Yuchen, et al.
Published: (2023)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
by: Jiang, Longtao, et al.
Published: (2025)
by: Jiang, Longtao, et al.
Published: (2025)
Similar Items
-
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
by: Zhu, Yuhan, et al.
Published: (2024) -
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
by: Zhang, Jiaming, et al.
Published: (2023) -
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024) -
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
by: Li, Kunchang, et al.
Published: (2023) -
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023)