CLIP Model for Images to Textual Prompts Based on Top-k Neighbors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xin, Cai, YeMing, Jia, Tianzhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
von: Yi, Chao, et al.
Veröffentlicht: (2024)
von: Yi, Chao, et al.
Veröffentlicht: (2024)
FG-CLIP: Fine-Grained Visual and Textual Alignment
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
von: Liu, Zichuan, et al.
Veröffentlicht: (2025)
von: Liu, Zichuan, et al.
Veröffentlicht: (2025)
Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks
von: An, Yanru, et al.
Veröffentlicht: (2025)
von: An, Yanru, et al.
Veröffentlicht: (2025)
Enhancing Spatial Reasoning through Visual and Textual Thinking
von: Liang, Xun, et al.
Veröffentlicht: (2025)
von: Liang, Xun, et al.
Veröffentlicht: (2025)
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
Top-Down Semantic Refinement for Image Captioning
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Fairness Analysis of CLIP-Based Foundation Models for X-Ray Image Classification
von: Sun, Xiangyu, et al.
Veröffentlicht: (2025)
von: Sun, Xiangyu, et al.
Veröffentlicht: (2025)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
Knowledge-Base based Semantic Image Transmission Using CLIP
von: Li, Chongyang, et al.
Veröffentlicht: (2025)
von: Li, Chongyang, et al.
Veröffentlicht: (2025)
Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
von: Feng, Zhanbo, et al.
Veröffentlicht: (2023)
von: Feng, Zhanbo, et al.
Veröffentlicht: (2023)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
von: Jin, Xin, et al.
Veröffentlicht: (2026)
von: Jin, Xin, et al.
Veröffentlicht: (2026)
Tuning Vision-Language Models with Candidate Labels by Prompt Alignment
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning
von: Dimitrovski, Ivica, et al.
Veröffentlicht: (2025)
von: Dimitrovski, Ivica, et al.
Veröffentlicht: (2025)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
Explicit Uncertainty Modeling for Active CLIP Adaptation with Dual Prompt Tuning
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
von: Wang, Xiao, et al.
Veröffentlicht: (2023)
von: Wang, Xiao, et al.
Veröffentlicht: (2023)
FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot Learning
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
von: Gao, Bin-Bin, et al.
Veröffentlicht: (2025)
von: Gao, Bin-Bin, et al.
Veröffentlicht: (2025)
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization
von: Wang, Jianpeng, et al.
Veröffentlicht: (2026)
von: Wang, Jianpeng, et al.
Veröffentlicht: (2026)
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
von: Cai, Yuliang, et al.
Veröffentlicht: (2025)
von: Cai, Yuliang, et al.
Veröffentlicht: (2025)
VisionCLIP: An Med-AIGC based Ethical Language-Image Foundation Model for Generalizable Retina Image Analysis
von: Wei, Hao, et al.
Veröffentlicht: (2024)
von: Wei, Hao, et al.
Veröffentlicht: (2024)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
von: Kim, Jeongho, et al.
Veröffentlicht: (2024)
von: Kim, Jeongho, et al.
Veröffentlicht: (2024)
Adversarial Prompt Tuning for Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
von: Che, Chang, et al.
Veröffentlicht: (2024)
von: Che, Chang, et al.
Veröffentlicht: (2024)
Interpreting CLIP's Image Representation via Text-Based Decomposition
von: Gandelsman, Yossi, et al.
Veröffentlicht: (2023)
von: Gandelsman, Yossi, et al.
Veröffentlicht: (2023)
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2023)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2023)
Curriculum Prompting Foundation Models for Medical Image Segmentation
von: Zheng, Xiuqi, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuqi, et al.
Veröffentlicht: (2024)
MINT: Memory-Infused Prompt Tuning at Test-time for CLIP
von: Yi, Jiaming, et al.
Veröffentlicht: (2025)
von: Yi, Jiaming, et al.
Veröffentlicht: (2025)
Synthesize Privacy-Preserving High-Resolution Images via Private Textual Intermediaries
von: Wang, Haoxiang, et al.
Veröffentlicht: (2025)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2025)
Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner
von: Zhou, Yitong, et al.
Veröffentlicht: (2024)
von: Zhou, Yitong, et al.
Veröffentlicht: (2024)
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
Shiva-DiT: Residual-Based Differentiable Top-$k$ Selection for Efficient Diffusion Transformers
von: Zhang, Jiaji, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaji, et al.
Veröffentlicht: (2026)
BiPrompt: Bilateral Prompt Optimization for Visual and Textual Debiasing in Vision-Language Models
von: Gupta, Sunny, et al.
Veröffentlicht: (2026)
von: Gupta, Sunny, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
von: Yi, Chao, et al.
Veröffentlicht: (2024) -
FG-CLIP: Fine-Grained Visual and Textual Alignment
von: Xie, Chunyu, et al.
Veröffentlicht: (2025) -
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
von: Liu, Zichuan, et al.
Veröffentlicht: (2025) -
Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks
von: An, Yanru, et al.
Veröffentlicht: (2025) -
Enhancing Spatial Reasoning through Visual and Textual Thinking
von: Liang, Xun, et al.
Veröffentlicht: (2025)