Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Donggeun, Jo, Yujin, Lee, Myungjoo, Kim, Taesup |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
di: Kim, Donggeun, et al.
Pubblicazione: (2024)
di: Kim, Donggeun, et al.
Pubblicazione: (2024)
MAFA: Managing False Negatives for Vision-Language Pre-training
di: Byun, Jaeseok, et al.
Pubblicazione: (2023)
di: Byun, Jaeseok, et al.
Pubblicazione: (2023)
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
di: Jo, Yujin, et al.
Pubblicazione: (2026)
di: Jo, Yujin, et al.
Pubblicazione: (2026)
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
di: Tang, Longxiang, et al.
Pubblicazione: (2024)
di: Tang, Longxiang, et al.
Pubblicazione: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
Leveraging Prior Knowledge of Diffusion Model for Person Search
di: Kim, Giyeol, et al.
Pubblicazione: (2025)
di: Kim, Giyeol, et al.
Pubblicazione: (2025)
Angular Gradient Sign Method: Uncovering Vulnerabilities in Hyperbolic Networks
di: Jo, Minsoo, et al.
Pubblicazione: (2025)
di: Jo, Minsoo, et al.
Pubblicazione: (2025)
Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
di: Yang, Seoyun, et al.
Pubblicazione: (2025)
di: Yang, Seoyun, et al.
Pubblicazione: (2025)
Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
di: Kim, Gihoon, et al.
Pubblicazione: (2025)
di: Kim, Gihoon, et al.
Pubblicazione: (2025)
Pushing the Limits of Vision-Language Models in Remote Sensing without Human Annotations
di: Cha, Keumgang, et al.
Pubblicazione: (2024)
di: Cha, Keumgang, et al.
Pubblicazione: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
di: Gao, Yuan, et al.
Pubblicazione: (2024)
di: Gao, Yuan, et al.
Pubblicazione: (2024)
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
di: Shin, Joonghyuk, et al.
Pubblicazione: (2025)
di: Shin, Joonghyuk, et al.
Pubblicazione: (2025)
When Model Knowledge meets Diffusion Model: Diffusion-assisted Data-free Image Synthesis with Alignment of Domain and Class
di: Kim, Yujin, et al.
Pubblicazione: (2025)
di: Kim, Yujin, et al.
Pubblicazione: (2025)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
di: Zhang, Dingwen, et al.
Pubblicazione: (2024)
di: Zhang, Dingwen, et al.
Pubblicazione: (2024)
Active Prompt Learning with Vision-Language Model Priors
di: Kim, Hoyoung, et al.
Pubblicazione: (2024)
di: Kim, Hoyoung, et al.
Pubblicazione: (2024)
Debiasing Classifiers by Amplifying Bias with Latent Diffusion and Large Language Models
di: Ko, Donggeun, et al.
Pubblicazione: (2024)
di: Ko, Donggeun, et al.
Pubblicazione: (2024)
AAPL: Adding Attributes to Prompt Learning for Vision-Language Models
di: Kim, Gahyeon, et al.
Pubblicazione: (2024)
di: Kim, Gahyeon, et al.
Pubblicazione: (2024)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
di: Kim, Gahyeon, et al.
Pubblicazione: (2025)
di: Kim, Gahyeon, et al.
Pubblicazione: (2025)
Grounded Knowledge-Enhanced Medical Vision-Language Pre-training for Chest X-Ray
di: Deng, Qiao, et al.
Pubblicazione: (2024)
di: Deng, Qiao, et al.
Pubblicazione: (2024)
From Tokens to Photons: Test-Time Physical Prompting for Vision-Language Models
di: Im, Boyeong, et al.
Pubblicazione: (2025)
di: Im, Boyeong, et al.
Pubblicazione: (2025)
Pseudo-Prompt Generating in Pre-trained Vision-Language Models for Multi-Label Medical Image Classification
di: Ye, Yaoqin, et al.
Pubblicazione: (2024)
di: Ye, Yaoqin, et al.
Pubblicazione: (2024)
CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
di: Lin, Xiao, et al.
Pubblicazione: (2024)
di: Lin, Xiao, et al.
Pubblicazione: (2024)
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
di: Hua, Hang, et al.
Pubblicazione: (2024)
di: Hua, Hang, et al.
Pubblicazione: (2024)
DiffInject: Revisiting Debias via Synthetic Data Generation using Diffusion-based Style Injection
di: Ko, Donggeun, et al.
Pubblicazione: (2024)
di: Ko, Donggeun, et al.
Pubblicazione: (2024)
Patch-Level Kernel Alignment for Dense Self-Supervised Learning
di: Yeo, Juan, et al.
Pubblicazione: (2025)
di: Yeo, Juan, et al.
Pubblicazione: (2025)
PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
di: Seo, Wonyong, et al.
Pubblicazione: (2026)
di: Seo, Wonyong, et al.
Pubblicazione: (2026)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
di: Miao, Yongzhu, et al.
Pubblicazione: (2023)
di: Miao, Yongzhu, et al.
Pubblicazione: (2023)
P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models
di: Jung, Geunyoung, et al.
Pubblicazione: (2026)
di: Jung, Geunyoung, et al.
Pubblicazione: (2026)
Lightweight Model Pre-training via Language Guided Knowledge Distillation
di: Li, Mingsheng, et al.
Pubblicazione: (2024)
di: Li, Mingsheng, et al.
Pubblicazione: (2024)
Long-term Pre-training for Temporal Action Detection with Transformers
di: Kim, Jihwan, et al.
Pubblicazione: (2024)
di: Kim, Jihwan, et al.
Pubblicazione: (2024)
PointT2I: LLM-based text-to-image generation via keypoints
di: Lee, Taekyung, et al.
Pubblicazione: (2025)
di: Lee, Taekyung, et al.
Pubblicazione: (2025)
Sample-agnostic Adversarial Perturbation for Vision-Language Pre-training Models
di: Zheng, Haonan, et al.
Pubblicazione: (2024)
di: Zheng, Haonan, et al.
Pubblicazione: (2024)
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
di: Huang, Jiale, et al.
Pubblicazione: (2024)
di: Huang, Jiale, et al.
Pubblicazione: (2024)
DoMIX: An Efficient Framework for Exploiting Domain Knowledge in Fine-Tuning
di: Kim, Dohoon, et al.
Pubblicazione: (2025)
di: Kim, Dohoon, et al.
Pubblicazione: (2025)
Object-Centric World Model for Language-Guided Manipulation
di: Jeong, Youngjoon, et al.
Pubblicazione: (2025)
di: Jeong, Youngjoon, et al.
Pubblicazione: (2025)
Efficient Vision-Language Pre-training by Cluster Masking
di: Wei, Zihao, et al.
Pubblicazione: (2024)
di: Wei, Zihao, et al.
Pubblicazione: (2024)
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
di: Yeo, Juan, et al.
Pubblicazione: (2025)
di: Yeo, Juan, et al.
Pubblicazione: (2025)
Continual Forgetting for Pre-trained Vision Models
di: Zhao, Hongbo, et al.
Pubblicazione: (2024)
di: Zhao, Hongbo, et al.
Pubblicazione: (2024)
Spatial-and-Frequency-aware Restoration method for Images based on Diffusion Models
di: Lee, Kyungsung, et al.
Pubblicazione: (2024)
di: Lee, Kyungsung, et al.
Pubblicazione: (2024)
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
di: Mahdizadeh, Ailar, et al.
Pubblicazione: (2025)
di: Mahdizadeh, Ailar, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
di: Kim, Donggeun, et al.
Pubblicazione: (2024) -
MAFA: Managing False Negatives for Vision-Language Pre-training
di: Byun, Jaeseok, et al.
Pubblicazione: (2023) -
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
di: Jo, Yujin, et al.
Pubblicazione: (2026) -
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
di: Tang, Longxiang, et al.
Pubblicazione: (2024) -
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
di: Ma, Shuailei, et al.
Pubblicazione: (2023)