ComCLIP: Training-Free Compositional Image and Text Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Kenan, He, Xuehai, Xu, Ruize, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
Bridging the Gap Between Multimodal Foundation Models and World Models
von: He, Xuehai
Veröffentlicht: (2025)
von: He, Xuehai
Veröffentlicht: (2025)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
GRIT: Teaching MLLMs to Think with Images
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
von: He, Xuehai, et al.
Veröffentlicht: (2025)
von: He, Xuehai, et al.
Veröffentlicht: (2025)
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
von: Allgeuer, Philipp, et al.
Veröffentlicht: (2024)
von: Allgeuer, Philipp, et al.
Veröffentlicht: (2024)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
von: Asokan, Mothilal, et al.
Veröffentlicht: (2025)
von: Asokan, Mothilal, et al.
Veröffentlicht: (2025)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
von: Liu, Delong, et al.
Veröffentlicht: (2023)
von: Liu, Delong, et al.
Veröffentlicht: (2023)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
MobileCLIP2: Improving Multi-Modal Reinforced Training
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment
von: Jiang, Liyao, et al.
Veröffentlicht: (2026)
von: Jiang, Liyao, et al.
Veröffentlicht: (2026)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
von: He, Xuehai, et al.
Veröffentlicht: (2024)
von: He, Xuehai, et al.
Veröffentlicht: (2024)
Text-only Synthesis for Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
Auditing Gender Presentation Differences in Text-to-Image Models
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Match & Choose: Model Selection Framework for Fine-tuning Text-to-Image Diffusion Models
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
Teaching Text-to-Image Models to Communicate in Dialog
von: Sun, Xiaowen, et al.
Veröffentlicht: (2023)
von: Sun, Xiaowen, et al.
Veröffentlicht: (2023)
Image-Text Relation Prediction for Multilingual Tweets
von: Rikters, Matīss, et al.
Veröffentlicht: (2025)
von: Rikters, Matīss, et al.
Veröffentlicht: (2025)
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
von: Kwon, Jihoon, et al.
Veröffentlicht: (2025)
von: Kwon, Jihoon, et al.
Veröffentlicht: (2025)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
von: Mehta, Sachin, et al.
Veröffentlicht: (2024)
von: Mehta, Sachin, et al.
Veröffentlicht: (2024)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
von: Chen, Wenting, et al.
Veröffentlicht: (2023)
von: Chen, Wenting, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024) -
Bridging the Gap Between Multimodal Foundation Models and World Models
von: He, Xuehai
Veröffentlicht: (2025) -
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
von: Lewis, Martha, et al.
Veröffentlicht: (2022) -
GRIT: Teaching MLLMs to Think with Images
von: Fan, Yue, et al.
Veröffentlicht: (2025) -
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)