Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Chen, Seto, Skyler, Abnar, Samira, Grangier, David, Jaitly, Navdeep, Susskind, Josh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Matryoshka Diffusion Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2023)
von: Gu, Jiatao, et al.
Veröffentlicht: (2023)
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
Normalizing Flows are Capable Generative Models
von: Zhai, Shuangfei, et al.
Veröffentlicht: (2024)
von: Zhai, Shuangfei, et al.
Veröffentlicht: (2024)
Improving GFlowNets for Text-to-Image Diffusion Alignment
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024)
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
How PARTs assemble into wholes: Learning the relative composition of images
von: Ayoughi, Melika, et al.
Veröffentlicht: (2025)
von: Ayoughi, Melika, et al.
Veröffentlicht: (2025)
HyperCLIP: Adapting Vision-Language models with Hypernetworks
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
von: Huang, Chen, et al.
Veröffentlicht: (2026)
von: Huang, Chen, et al.
Veröffentlicht: (2026)
Normalizing Trajectory Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2026)
von: Gu, Jiatao, et al.
Veröffentlicht: (2026)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
von: Sampaio, Georgia Gabriela, et al.
Veröffentlicht: (2024)
von: Sampaio, Georgia Gabriela, et al.
Veröffentlicht: (2024)
Adapting to Distribution Shift by Visual Domain Prompt Generation
von: Chi, Zhixiang, et al.
Veröffentlicht: (2024)
von: Chi, Zhixiang, et al.
Veröffentlicht: (2024)
Enhancing CLIP with CLIP: Exploring Pseudolabeling for Limited-Label Prompt Tuning
von: Menghini, Cristina, et al.
Veröffentlicht: (2023)
von: Menghini, Cristina, et al.
Veröffentlicht: (2023)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
von: Gu, Jiatao, et al.
Veröffentlicht: (2025)
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
von: Li, Xianhang, et al.
Veröffentlicht: (2025)
Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performance
von: Guarino, Vanessa Emanuela, et al.
Veröffentlicht: (2026)
von: Guarino, Vanessa Emanuela, et al.
Veröffentlicht: (2026)
Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain Adaptation
von: Chi, Zhixiang, et al.
Veröffentlicht: (2025)
von: Chi, Zhixiang, et al.
Veröffentlicht: (2025)
MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental Learning
von: Nicolas, Julien, et al.
Veröffentlicht: (2023)
von: Nicolas, Julien, et al.
Veröffentlicht: (2023)
How Far Are We from Intelligent Visual Deductive Reasoning?
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent Modeling
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
von: Gu, Jiatao, et al.
Veröffentlicht: (2024)
The Coupling Within: Flow Matching via Distilled Normalizing Flows
von: Berthelot, David, et al.
Veröffentlicht: (2026)
von: Berthelot, David, et al.
Veröffentlicht: (2026)
Understanding Model Reprogramming for CLIP via Decoupling Visual Prompts
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
von: Cai, Chengyi, et al.
Veröffentlicht: (2025)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
von: Medeiros, Heitor R., et al.
Veröffentlicht: (2024)
von: Medeiros, Heitor R., et al.
Veröffentlicht: (2024)
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks
von: Ding, Yuhe, et al.
Veröffentlicht: (2024)
von: Ding, Yuhe, et al.
Veröffentlicht: (2024)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
von: Grangier, David, et al.
Veröffentlicht: (2024)
von: Grangier, David, et al.
Veröffentlicht: (2024)
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
von: Shen, Ying, et al.
Veröffentlicht: (2026)
von: Shen, Ying, et al.
Veröffentlicht: (2026)
Federated Domain Generalization via Prompt Learning and Aggregation
von: Gong, Shuai, et al.
Veröffentlicht: (2024)
von: Gong, Shuai, et al.
Veröffentlicht: (2024)
Transitive Vision-Language Prompt Learning for Domain Generalization
von: Wang, Liyuan, et al.
Veröffentlicht: (2024)
von: Wang, Liyuan, et al.
Veröffentlicht: (2024)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
von: Zang, Yuhang, et al.
Veröffentlicht: (2024)
von: Zang, Yuhang, et al.
Veröffentlicht: (2024)
MIP: CLIP-based Image Reconstruction from PEFT Gradients
von: Zhou, Peiheng, et al.
Veröffentlicht: (2024)
von: Zhou, Peiheng, et al.
Veröffentlicht: (2024)
Detecting AI-Generated Images via CLIP
von: Moskowitz, A. G., et al.
Veröffentlicht: (2024)
von: Moskowitz, A. G., et al.
Veröffentlicht: (2024)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
von: Wang, Xinze, et al.
Veröffentlicht: (2025)
von: Wang, Xinze, et al.
Veröffentlicht: (2025)
Learning Generalizable Prompt for CLIP with Class Similarity Knowledge
von: Jung, Sehun, et al.
Veröffentlicht: (2025)
von: Jung, Sehun, et al.
Veröffentlicht: (2025)
Breaking the Limits of Open-Weight CLIP: An Optimization Framework for Self-supervised Fine-tuning of CLIP
von: Mehta, Anant, et al.
Veröffentlicht: (2026)
von: Mehta, Anant, et al.
Veröffentlicht: (2026)
DeCLIP: Decoding CLIP representations for deepfake localization
von: Smeu, Stefan, et al.
Veröffentlicht: (2024)
von: Smeu, Stefan, et al.
Veröffentlicht: (2024)
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
von: Singha, Mainak, et al.
Veröffentlicht: (2023)
von: Singha, Mainak, et al.
Veröffentlicht: (2023)
FastCLIP: A Suite of Optimization Techniques to Accelerate CLIP Training with Limited Resources
von: Wei, Xiyuan, et al.
Veröffentlicht: (2024)
von: Wei, Xiyuan, et al.
Veröffentlicht: (2024)
Stable Diffusion Dataset Generation for Downstream Classification Tasks
von: Lomurno, Eugenio, et al.
Veröffentlicht: (2024)
von: Lomurno, Eugenio, et al.
Veröffentlicht: (2024)
When and How Does CLIP Enable Domain and Compositional Generalization?
von: Kempf, Elias, et al.
Veröffentlicht: (2025)
von: Kempf, Elias, et al.
Veröffentlicht: (2025)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
von: Magistri, Simone, et al.
Veröffentlicht: (2026)
von: Magistri, Simone, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Matryoshka Diffusion Models
von: Gu, Jiatao, et al.
Veröffentlicht: (2023) -
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
von: Gu, Jiatao, et al.
Veröffentlicht: (2024) -
Normalizing Flows are Capable Generative Models
von: Zhai, Shuangfei, et al.
Veröffentlicht: (2024) -
Improving GFlowNets for Text-to-Image Diffusion Alignment
von: Zhang, Dinghuai, et al.
Veröffentlicht: (2024) -
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
von: Huang, Chen, et al.
Veröffentlicht: (2025)