Asymmetric Masked Distillation for Pre-Training Small Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Zhiyu, Huang, Bingkun, Xing, Sen, Wu, Gangshan, Qiao, Yu, Wang, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
MixFormerV2: Efficient Fully Transformer Tracking
von: Cui, Yutao, et al.
Veröffentlicht: (2023)
von: Cui, Yutao, et al.
Veröffentlicht: (2023)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
Data-efficient Event Camera Pre-training via Disentangled Masked Modeling
von: Huang, Zhenpeng, et al.
Veröffentlicht: (2024)
von: Huang, Zhenpeng, et al.
Veröffentlicht: (2024)
STMixer: A One-Stage Sparse Action Detector
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Open-Vocabulary Spatio-Temporal Action Detection
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Dual DETRs for Multi-Label Temporal Action Detection
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
Pure-Pass: Fine-Grained, Adaptive Masking for Dynamic Token-Mixing Routing in Lightweight Image Super-Resolution
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
Efficient Test-Time Prompt Tuning for Vision-Language Models
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
von: Luo, Xingyu, et al.
Veröffentlicht: (2026)
von: Luo, Xingyu, et al.
Veröffentlicht: (2026)
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
von: Wu, Biao, et al.
Veröffentlicht: (2024)
von: Wu, Biao, et al.
Veröffentlicht: (2024)
Harvest Video Foundation Models via Efficient Post-Pretraining
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
MiM: Mask in Mask Self-Supervised Pre-Training for 3D Medical Image Analysis
von: Zhuang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhuang, Jiaxin, et al.
Veröffentlicht: (2024)
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
von: Zhu, Haoyi, et al.
Veröffentlicht: (2023)
von: Zhu, Haoyi, et al.
Veröffentlicht: (2023)
Pre-Trained Masked Image Model for Mobile Robot Navigation
von: Sharma, Vishnu Dutt, et al.
Veröffentlicht: (2023)
von: Sharma, Vishnu Dutt, et al.
Veröffentlicht: (2023)
Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
von: Rang, Miao, et al.
Veröffentlicht: (2025)
von: Rang, Miao, et al.
Veröffentlicht: (2025)
Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
von: Röhrich, Nikolai, et al.
Veröffentlicht: (2025)
von: Röhrich, Nikolai, et al.
Veröffentlicht: (2025)
Multi-Modal Soccer Scene Analysis with Masked Pre-Training
von: Peral, Marc, et al.
Veröffentlicht: (2025)
von: Peral, Marc, et al.
Veröffentlicht: (2025)
SOEDiff: Efficient Distillation for Small Object Editing
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
An Experimental Study on Exploring Strong Lightweight Vision Transformers via Masked Image Modeling Pre-Training
von: Gao, Jin, et al.
Veröffentlicht: (2024)
von: Gao, Jin, et al.
Veröffentlicht: (2024)
CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
von: Wu, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoxue, et al.
Veröffentlicht: (2025)
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models
von: Fan, Jiawei, et al.
Veröffentlicht: (2026)
von: Fan, Jiawei, et al.
Veröffentlicht: (2026)
ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal Prediction
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2023)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2023)
Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching
von: Xia, Qianxin, et al.
Veröffentlicht: (2026)
von: Xia, Qianxin, et al.
Veröffentlicht: (2026)
ComKD-CLIP: Comprehensive Knowledge Distillation for Contrastive Language-Image Pre-traning Model
von: Chen, Yifan, et al.
Veröffentlicht: (2024)
von: Chen, Yifan, et al.
Veröffentlicht: (2024)
Reprogramming Distillation for Medical Foundation Models
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
Lightweight Model Pre-training via Language Guided Knowledge Distillation
von: Li, Mingsheng, et al.
Veröffentlicht: (2024)
von: Li, Mingsheng, et al.
Veröffentlicht: (2024)
Deep Reprogramming Distillation for Medical Foundation Models
von: Du, Siyuan, et al.
Veröffentlicht: (2026)
von: Du, Siyuan, et al.
Veröffentlicht: (2026)
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
von: Chen, Yitong, et al.
Veröffentlicht: (2025)
von: Chen, Yitong, et al.
Veröffentlicht: (2025)
Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation
von: Yang, Yuchen, et al.
Veröffentlicht: (2023)
von: Yang, Yuchen, et al.
Veröffentlicht: (2023)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
von: Jiang, Longtao, et al.
Veröffentlicht: (2025)
von: Jiang, Longtao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024) -
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023) -
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
von: Wu, Tao, et al.
Veröffentlicht: (2024) -
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
von: Li, Kunchang, et al.
Veröffentlicht: (2023) -
MixFormerV2: Efficient Fully Transformer Tracking
von: Cui, Yutao, et al.
Veröffentlicht: (2023)