MoDE: CLIP Data Experts via Clustering
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Jiawei, Huang, Po-Yao, Xie, Saining, Li, Shang-Wen, Zettlemoyer, Luke, Chang, Shih-Fu, Yih, Wen-Tau, Xu, Hu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demystifying CLIP Data
by: Xu, Hu, et al.
Published: (2023)
by: Xu, Hu, et al.
Published: (2023)
Altogether: Image Captioning via Re-aligning Alt-text
by: Xu, Hu, et al.
Published: (2024)
by: Xu, Hu, et al.
Published: (2024)
Meta CLIP 2: A Worldwide Scaling Recipe
by: Chuang, Yung-Sung, et al.
Published: (2025)
by: Chuang, Yung-Sung, et al.
Published: (2025)
Continual Learning via Sparse Memory Finetuning
by: Lin, Jessy, et al.
Published: (2025)
by: Lin, Jessy, et al.
Published: (2025)
MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
by: Fan, Qiannan, et al.
Published: (2025)
by: Fan, Qiannan, et al.
Published: (2025)
MoDE: Effective Multi-task Parameter Efficient Fine-Tuning with a Mixture of Dyadic Experts
by: Ning, Lin, et al.
Published: (2024)
by: Ning, Lin, et al.
Published: (2024)
Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model
by: He, Jacqueline, et al.
Published: (2026)
by: He, Jacqueline, et al.
Published: (2026)
Memory Layers at Scale
by: Berges, Vincent-Pierre, et al.
Published: (2024)
by: Berges, Vincent-Pierre, et al.
Published: (2024)
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
by: Xie, Zhitian, et al.
Published: (2024)
by: Xie, Zhitian, et al.
Published: (2024)
Slicing and Dicing: Configuring Optimal Mixtures of Experts
by: Li, Margaret, et al.
Published: (2026)
by: Li, Margaret, et al.
Published: (2026)
Improving Factuality with Explicit Working Memory
by: Chen, Mingda, et al.
Published: (2024)
by: Chen, Mingda, et al.
Published: (2024)
Reliable, Adaptable, and Attributable Language Models with Retrieval
by: Asai, Akari, et al.
Published: (2024)
by: Asai, Akari, et al.
Published: (2024)
Text Quality-Based Pruning for Efficient Training of Language Models
by: Sharma, Vasu, et al.
Published: (2024)
by: Sharma, Vasu, et al.
Published: (2024)
Learning Facts at Scale with Active Reading
by: Lin, Jessy, et al.
Published: (2025)
by: Lin, Jessy, et al.
Published: (2025)
MoExtend: Tuning New Experts for Modality and Task Extension
by: Zhong, Shanshan, et al.
Published: (2024)
by: Zhong, Shanshan, et al.
Published: (2024)
Few-Shot Data Synthesis for Open Domain Multi-Hop Question Answering
by: Chen, Mingda, et al.
Published: (2023)
by: Chen, Mingda, et al.
Published: (2023)
MatchNAS: Optimizing Edge AI in Sparse-Label Data Contexts via Automating Deep Neural Network Porting for Mobile Deployment
by: Huang, Hongtao, et al.
Published: (2024)
by: Huang, Hongtao, et al.
Published: (2024)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
Flow Map Distillation Without Data
by: Tong, Shangyuan, et al.
Published: (2025)
by: Tong, Shangyuan, et al.
Published: (2025)
Dual-disentangled Deep Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)
by: Yao, Jiawei, et al.
Published: (2024)
Micro Language Models Enable Instant Responses
by: Cheng, Wen, et al.
Published: (2026)
by: Cheng, Wen, et al.
Published: (2026)
Breaking the Curse of Multilinguality with Cross-lingual Expert Language Models
by: Blevins, Terra, et al.
Published: (2024)
by: Blevins, Terra, et al.
Published: (2024)
WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization
by: Ma, Jiawei, et al.
Published: (2024)
by: Ma, Jiawei, et al.
Published: (2024)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
by: Wang, Hsuan-Fu, et al.
Published: (2024)
by: Wang, Hsuan-Fu, et al.
Published: (2024)
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
MH-MoE: Multi-Head Mixture-of-Experts
by: Huang, Shaohan, et al.
Published: (2024)
by: Huang, Shaohan, et al.
Published: (2024)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
Comparing Hallucination Detection Metrics for Multilingual Generation
by: Kang, Haoqiang, et al.
Published: (2024)
by: Kang, Haoqiang, et al.
Published: (2024)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
by: Zhang, Kangjie, et al.
Published: (2026)
by: Zhang, Kangjie, et al.
Published: (2026)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
by: Yao, Xin, et al.
Published: (2025)
by: Yao, Xin, et al.
Published: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
by: Sukhbaatar, Sainbayar, et al.
Published: (2024)
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
by: Zhu, Peijun, et al.
Published: (2025)
by: Zhu, Peijun, et al.
Published: (2025)
Demystifying Prompts in Language Models via Perplexity Estimation
by: Gonen, Hila, et al.
Published: (2022)
by: Gonen, Hila, et al.
Published: (2022)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
by: Huang, Yushi, et al.
Published: (2025)
by: Huang, Yushi, et al.
Published: (2025)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025)
by: Yasunaga, Michihiro, et al.
Published: (2025)
NeuroCLIP: Neuromorphic Data Understanding by CLIP and SNN
by: Guo, Yufei, et al.
Published: (2023)
by: Guo, Yufei, et al.
Published: (2023)
Similar Items
-
Demystifying CLIP Data
by: Xu, Hu, et al.
Published: (2023) -
Altogether: Image Captioning via Re-aligning Alt-text
by: Xu, Hu, et al.
Published: (2024) -
Meta CLIP 2: A Worldwide Scaling Recipe
by: Chuang, Yung-Sung, et al.
Published: (2025) -
Continual Learning via Sparse Memory Finetuning
by: Lin, Jessy, et al.
Published: (2025) -
MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
by: Fan, Qiannan, et al.
Published: (2025)