MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hui, Yang, Pengfei, Chen, Juanyang, Dong, Le, Chen, Yanxin, Wang, Quan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
Knowledge Bridger: Towards Training-free Missing Modality Completion
by: Ke, Guanzhou, et al.
Published: (2025)
by: Ke, Guanzhou, et al.
Published: (2025)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
by: Zhang, Kangkai, et al.
Published: (2024)
by: Zhang, Kangkai, et al.
Published: (2024)
Cross-Modal Coordination Across a Diverse Set of Input Modalities
by: Sánchez, Jorge, et al.
Published: (2024)
by: Sánchez, Jorge, et al.
Published: (2024)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Overcome Modal Bias in Multi-modal Federated Learning via Balanced Modality Selection
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
by: Zhang, Yafei, et al.
Published: (2025)
by: Zhang, Yafei, et al.
Published: (2025)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
by: Luo, Qiuming, et al.
Published: (2026)
by: Luo, Qiuming, et al.
Published: (2026)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention
by: Panta, Sanjeev, et al.
Published: (2026)
by: Panta, Sanjeev, et al.
Published: (2026)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
by: Wang, Yabing, et al.
Published: (2023)
by: Wang, Yabing, et al.
Published: (2023)
Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
by: Shi, Ruixin, et al.
Published: (2024)
by: Shi, Ruixin, et al.
Published: (2024)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
by: Saleh, Mohamed, et al.
Published: (2026)
by: Saleh, Mohamed, et al.
Published: (2026)
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
by: Luo, Bingjun, et al.
Published: (2025)
by: Luo, Bingjun, et al.
Published: (2025)
Knowledge Distillation Based on Transformed Teacher Matching
by: Zheng, Kaixiang, et al.
Published: (2024)
by: Zheng, Kaixiang, et al.
Published: (2024)
Identity Preserving 3D Head Stylization with Multiview Score Distillation
by: Bilecen, Bahri Batuhan, et al.
Published: (2024)
by: Bilecen, Bahri Batuhan, et al.
Published: (2024)
BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
by: Zhang, Zheng, et al.
Published: (2024)
by: Zhang, Zheng, et al.
Published: (2024)
MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates
by: Zhao, Binyu, et al.
Published: (2025)
by: Zhao, Binyu, et al.
Published: (2025)
Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
by: Lin, Yan-Bo, et al.
Published: (2025)
by: Lin, Yan-Bo, et al.
Published: (2025)
Self-similarity Prior Distillation for Unsupervised Remote Physiological Measurement
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Competitive Learning for Achieving Content-specific Filters in Video Coding for Machines
by: Zhang, Honglei, et al.
Published: (2024)
by: Zhang, Honglei, et al.
Published: (2024)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
FeatDistill: A Feature Distillation Enhanced Multi-Expert Ensemble Framework for Robust AI-generated Image Detection
by: Tu, Zhilin, et al.
Published: (2026)
by: Tu, Zhilin, et al.
Published: (2026)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
by: Huang, Shunyu, et al.
Published: (2026)
by: Huang, Shunyu, et al.
Published: (2026)
Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning
by: Wang, Jinpeng, et al.
Published: (2025)
by: Wang, Jinpeng, et al.
Published: (2025)
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts
by: Zhong, Guowei, et al.
Published: (2025)
by: Zhong, Guowei, et al.
Published: (2025)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
by: Lin, Xiang, et al.
Published: (2025)
by: Lin, Xiang, et al.
Published: (2025)
A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
by: Dai, Yusheng, et al.
Published: (2024)
by: Dai, Yusheng, et al.
Published: (2024)
ReconBoost: Boosting Can Achieve Modality Reconcilement
by: Hua, Cong, et al.
Published: (2024)
by: Hua, Cong, et al.
Published: (2024)
Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction
by: Mu, Zhaoxi, et al.
Published: (2024)
by: Mu, Zhaoxi, et al.
Published: (2024)
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
Rethinking Multi-view Representation Learning via Distilled Disentangling
by: Ke, Guanzhou, et al.
Published: (2024)
by: Ke, Guanzhou, et al.
Published: (2024)
DeCo-VAE: Learning Compact Latents for Video Reconstruction via Decoupled Representation
by: Yin, Xiangchen, et al.
Published: (2025)
by: Yin, Xiangchen, et al.
Published: (2025)
Similar Items
-
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
by: Chen, Yang, et al.
Published: (2024) -
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025) -
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
by: Mekonnen, Kidist Amde, et al.
Published: (2024) -
Knowledge Bridger: Towards Training-free Missing Modality Completion
by: Ke, Guanzhou, et al.
Published: (2025) -
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
by: Croitoru, Florinel-Alin, et al.
Published: (2022)