Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ruiming, Yang, Junming, Xia, Shiyu, Yang, Xu, Wang, Jing, Geng, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models
by: Xia, Shi-Yu, et al.
Published: (2024)
by: Xia, Shi-Yu, et al.
Published: (2024)
CLIP-DFGS: A Hard Sample Mining Method for CLIP in Generalizable Person Re-Identification
by: Zhao, Huazhong, et al.
Published: (2024)
by: Zhao, Huazhong, et al.
Published: (2024)
BadCLIP++: Stealthy and Persistent Backdoors in Multimodal Contrastive Learning
by: Liang, Siyuan, et al.
Published: (2026)
by: Liang, Siyuan, et al.
Published: (2026)
Multimodal Multilabel Classification by CLIP
by: Guo, Yanming
Published: (2024)
by: Guo, Yanming
Published: (2024)
DiffCLIP: Few-shot Language-driven Multimodal Classifier
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks
by: Jin, Jing, et al.
Published: (2026)
by: Jin, Jing, et al.
Published: (2026)
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
by: Chen, Huiyi, et al.
Published: (2025)
by: Chen, Huiyi, et al.
Published: (2025)
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition
by: Xing, Jiazheng, et al.
Published: (2023)
by: Xing, Jiazheng, et al.
Published: (2023)
EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation
by: Zhang, Zelin, et al.
Published: (2025)
by: Zhang, Zelin, et al.
Published: (2025)
A CLIP-Powered Framework for Robust and Generalizable Data Selection
by: Yang, Suorong, et al.
Published: (2024)
by: Yang, Suorong, et al.
Published: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
by: Song, Yehun, et al.
Published: (2025)
by: Song, Yehun, et al.
Published: (2025)
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
by: Wang, Zhu, et al.
Published: (2025)
by: Wang, Zhu, et al.
Published: (2025)
Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
by: Mei, Guofeng, et al.
Published: (2025)
by: Mei, Guofeng, et al.
Published: (2025)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
by: Chung, Jeannie, et al.
Published: (2026)
by: Chung, Jeannie, et al.
Published: (2026)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
by: Jiang, Tianxiang, et al.
Published: (2025)
by: Jiang, Tianxiang, et al.
Published: (2025)
Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
by: Liu, Huimin, et al.
Published: (2025)
by: Liu, Huimin, et al.
Published: (2025)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
by: Ali, Muhammad, et al.
Published: (2024)
by: Ali, Muhammad, et al.
Published: (2024)
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
by: Zhang, Yuchen, et al.
Published: (2026)
by: Zhang, Yuchen, et al.
Published: (2026)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
by: Yang, Shengzhu, et al.
Published: (2025)
by: Yang, Shengzhu, et al.
Published: (2025)
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
by: Zhang, Daoan, et al.
Published: (2024)
by: Zhang, Daoan, et al.
Published: (2024)
UniCrossAdapter: Multimodal Adaptation of CLIP for Radiology Report Generation
by: Chen, Yaxiong, et al.
Published: (2025)
by: Chen, Yaxiong, et al.
Published: (2025)
GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
by: Wang, Jun, et al.
Published: (2023)
by: Wang, Jun, et al.
Published: (2023)
DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
by: Yang, Zhiwei, et al.
Published: (2026)
by: Yang, Zhiwei, et al.
Published: (2026)
LLM-driven Knowledge Enhancement for Multimodal Cancer Survival Prediction
by: Zhao, Chenyu, et al.
Published: (2025)
by: Zhao, Chenyu, et al.
Published: (2025)
MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
by: Liu, Yepeng, et al.
Published: (2025)
by: Liu, Yepeng, et al.
Published: (2025)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2025)
by: Xiong, Zhitong, et al.
Published: (2025)
PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest
by: Beal, Josh, et al.
Published: (2026)
by: Beal, Josh, et al.
Published: (2026)
Multimodal CLIP Inference for Meta-Few-Shot Image Classification
by: Ferragu, Constance, et al.
Published: (2024)
by: Ferragu, Constance, et al.
Published: (2024)
GatedCLIP: Gated Multimodal Fusion for Hateful Memes Detection
by: Guo, Yingying, et al.
Published: (2026)
by: Guo, Yingying, et al.
Published: (2026)
Generalizable Prompt Learning of CLIP: A Brief Overview
by: Cui, Fangming, et al.
Published: (2025)
by: Cui, Fangming, et al.
Published: (2025)
Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities
by: Li, Mingcheng, et al.
Published: (2024)
by: Li, Mingcheng, et al.
Published: (2024)
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection
by: Qin, Haotian, et al.
Published: (2025)
by: Qin, Haotian, et al.
Published: (2025)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
by: Wang, Mengmeng, et al.
Published: (2024)
by: Wang, Mengmeng, et al.
Published: (2024)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
by: Luo, Meng, et al.
Published: (2026)
by: Luo, Meng, et al.
Published: (2026)
Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
by: Liang, Xusheng, et al.
Published: (2025)
by: Liang, Xusheng, et al.
Published: (2025)
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Learning Generalizable Prompt for CLIP with Class Similarity Knowledge
by: Jung, Sehun, et al.
Published: (2025)
by: Jung, Sehun, et al.
Published: (2025)
Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
by: Xu, Ganxi, et al.
Published: (2025)
by: Xu, Ganxi, et al.
Published: (2025)
Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection
by: Yermakov, Andrii, et al.
Published: (2025)
by: Yermakov, Andrii, et al.
Published: (2025)
Interactive Multimodal Fusion with Temporal Modeling
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
Similar Items
-
Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models
by: Xia, Shi-Yu, et al.
Published: (2024) -
CLIP-DFGS: A Hard Sample Mining Method for CLIP in Generalizable Person Re-Identification
by: Zhao, Huazhong, et al.
Published: (2024) -
BadCLIP++: Stealthy and Persistent Backdoors in Multimodal Contrastive Learning
by: Liang, Siyuan, et al.
Published: (2026) -
Multimodal Multilabel Classification by CLIP
by: Guo, Yanming
Published: (2024) -
DiffCLIP: Few-shot Language-driven Multimodal Classifier
by: Zhang, Jiaqing, et al.
Published: (2024)