Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Chen, Seto, Skyler, Abnar, Samira, Grangier, David, Jaitly, Navdeep, Susskind, Josh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Matryoshka Diffusion Models
by: Gu, Jiatao, et al.
Published: (2023)
by: Gu, Jiatao, et al.
Published: (2023)
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
Normalizing Flows are Capable Generative Models
by: Zhai, Shuangfei, et al.
Published: (2024)
by: Zhai, Shuangfei, et al.
Published: (2024)
Improving GFlowNets for Text-to-Image Diffusion Alignment
by: Zhang, Dinghuai, et al.
Published: (2024)
by: Zhang, Dinghuai, et al.
Published: (2024)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
by: Huang, Chen, et al.
Published: (2025)
by: Huang, Chen, et al.
Published: (2025)
How PARTs assemble into wholes: Learning the relative composition of images
by: Ayoughi, Melika, et al.
Published: (2025)
by: Ayoughi, Melika, et al.
Published: (2025)
HyperCLIP: Adapting Vision-Language models with Hypernetworks
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
Normalizing Trajectory Models
by: Gu, Jiatao, et al.
Published: (2026)
by: Gu, Jiatao, et al.
Published: (2026)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
Adapting to Distribution Shift by Visual Domain Prompt Generation
by: Chi, Zhixiang, et al.
Published: (2024)
by: Chi, Zhixiang, et al.
Published: (2024)
Enhancing CLIP with CLIP: Exploring Pseudolabeling for Limited-Label Prompt Tuning
by: Menghini, Cristina, et al.
Published: (2023)
by: Menghini, Cristina, et al.
Published: (2023)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
by: Li, Xianhang, et al.
Published: (2025)
by: Li, Xianhang, et al.
Published: (2025)
Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performance
by: Guarino, Vanessa Emanuela, et al.
Published: (2026)
by: Guarino, Vanessa Emanuela, et al.
Published: (2026)
Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain Adaptation
by: Chi, Zhixiang, et al.
Published: (2025)
by: Chi, Zhixiang, et al.
Published: (2025)
MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental Learning
by: Nicolas, Julien, et al.
Published: (2023)
by: Nicolas, Julien, et al.
Published: (2023)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning
by: Xie, Zhen-Hao, et al.
Published: (2026)
by: Xie, Zhen-Hao, et al.
Published: (2026)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
The Coupling Within: Flow Matching via Distilled Normalizing Flows
by: Berthelot, David, et al.
Published: (2026)
by: Berthelot, David, et al.
Published: (2026)
Understanding Model Reprogramming for CLIP via Decoupling Visual Prompts
by: Cai, Chengyi, et al.
Published: (2025)
by: Cai, Chengyi, et al.
Published: (2025)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks
by: Ding, Yuhe, et al.
Published: (2024)
by: Ding, Yuhe, et al.
Published: (2024)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
Federated Domain Generalization via Prompt Learning and Aggregation
by: Gong, Shuai, et al.
Published: (2024)
by: Gong, Shuai, et al.
Published: (2024)
Transitive Vision-Language Prompt Learning for Domain Generalization
by: Wang, Liyuan, et al.
Published: (2024)
by: Wang, Liyuan, et al.
Published: (2024)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
by: Zang, Yuhang, et al.
Published: (2024)
by: Zang, Yuhang, et al.
Published: (2024)
MIP: CLIP-based Image Reconstruction from PEFT Gradients
by: Zhou, Peiheng, et al.
Published: (2024)
by: Zhou, Peiheng, et al.
Published: (2024)
Detecting AI-Generated Images via CLIP
by: Moskowitz, A. G., et al.
Published: (2024)
by: Moskowitz, A. G., et al.
Published: (2024)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)
by: Wang, Xinze, et al.
Published: (2025)
Learning Generalizable Prompt for CLIP with Class Similarity Knowledge
by: Jung, Sehun, et al.
Published: (2025)
by: Jung, Sehun, et al.
Published: (2025)
Breaking the Limits of Open-Weight CLIP: An Optimization Framework for Self-supervised Fine-tuning of CLIP
by: Mehta, Anant, et al.
Published: (2026)
by: Mehta, Anant, et al.
Published: (2026)
DeCLIP: Decoding CLIP representations for deepfake localization
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
by: Singha, Mainak, et al.
Published: (2023)
by: Singha, Mainak, et al.
Published: (2023)
FastCLIP: A Suite of Optimization Techniques to Accelerate CLIP Training with Limited Resources
by: Wei, Xiyuan, et al.
Published: (2024)
by: Wei, Xiyuan, et al.
Published: (2024)
Stable Diffusion Dataset Generation for Downstream Classification Tasks
by: Lomurno, Eugenio, et al.
Published: (2024)
by: Lomurno, Eugenio, et al.
Published: (2024)
When and How Does CLIP Enable Domain and Compositional Generalization?
by: Kempf, Elias, et al.
Published: (2025)
by: Kempf, Elias, et al.
Published: (2025)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
by: Magistri, Simone, et al.
Published: (2026)
by: Magistri, Simone, et al.
Published: (2026)
Similar Items
-
Matryoshka Diffusion Models
by: Gu, Jiatao, et al.
Published: (2023) -
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
by: Gu, Jiatao, et al.
Published: (2024) -
Normalizing Flows are Capable Generative Models
by: Zhai, Shuangfei, et al.
Published: (2024) -
Improving GFlowNets for Text-to-Image Diffusion Alignment
by: Zhang, Dinghuai, et al.
Published: (2024) -
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
by: Huang, Chen, et al.
Published: (2025)