Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kwon, Jihoon, Min, Kyle, Sohn, Jy-yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
von: Kwon, Jihoon, et al.
Veröffentlicht: (2026)
von: Kwon, Jihoon, et al.
Veröffentlicht: (2026)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
von: Jeon, Jaebyeong, et al.
Veröffentlicht: (2025)
von: Jeon, Jaebyeong, et al.
Veröffentlicht: (2025)
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
von: Shivika, et al.
Veröffentlicht: (2026)
von: Shivika, et al.
Veröffentlicht: (2026)
IDEA: Image Description Enhanced CLIP-Adapter
von: Ye, Zhipeng, et al.
Veröffentlicht: (2025)
von: Ye, Zhipeng, et al.
Veröffentlicht: (2025)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2024)
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
von: Che, Chang, et al.
Veröffentlicht: (2024)
von: Che, Chang, et al.
Veröffentlicht: (2024)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
von: Asokan, Mothilal, et al.
Veröffentlicht: (2025)
von: Asokan, Mothilal, et al.
Veröffentlicht: (2025)
ComCLIP: Training-Free Compositional Image and Text Matching
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
A Theoretical Framework for Preventing Class Collapse in Supervised Contrastive Learning
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
Interpreting CLIP's Image Representation via Text-Based Decomposition
von: Gandelsman, Yossi, et al.
Veröffentlicht: (2023)
von: Gandelsman, Yossi, et al.
Veröffentlicht: (2023)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
von: Ma, Wenxin, et al.
Veröffentlicht: (2025)
von: Ma, Wenxin, et al.
Veröffentlicht: (2025)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2026)
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2026)
Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis
von: Samanta, Argha Kamal, et al.
Veröffentlicht: (2025)
von: Samanta, Argha Kamal, et al.
Veröffentlicht: (2025)
FG-CLIP: Fine-Grained Visual and Textual Alignment
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
von: Cao, Min, et al.
Veröffentlicht: (2025)
von: Cao, Min, et al.
Veröffentlicht: (2025)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
V-LynX: Token Interface Alignment for Video+X LLMs
von: Park, Jungin, et al.
Veröffentlicht: (2026)
von: Park, Jungin, et al.
Veröffentlicht: (2026)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
von: He, Chiyuan, et al.
Veröffentlicht: (2025)
von: He, Chiyuan, et al.
Veröffentlicht: (2025)
DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment
von: Chen, Weizhi, et al.
Veröffentlicht: (2025)
von: Chen, Weizhi, et al.
Veröffentlicht: (2025)
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
von: Xie, Shaoan, et al.
Veröffentlicht: (2025)
von: Xie, Shaoan, et al.
Veröffentlicht: (2025)
Enhancing Source-Free Domain Adaptive Object Detection with Low-confidence Pseudo Label Distillation
von: Yoon, Ilhoon, et al.
Veröffentlicht: (2024)
von: Yoon, Ilhoon, et al.
Veröffentlicht: (2024)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
von: Kim, Jiyeong, et al.
Veröffentlicht: (2026)
von: Kim, Jiyeong, et al.
Veröffentlicht: (2026)
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
von: Xue, Leyan, et al.
Veröffentlicht: (2025)
von: Xue, Leyan, et al.
Veröffentlicht: (2025)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
von: Yang, Shengzhu, et al.
Veröffentlicht: (2025)
von: Yang, Shengzhu, et al.
Veröffentlicht: (2025)
Discriminative Perception via Anchored Description for Reasoning Segmentation
von: Yang, Tao, et al.
Veröffentlicht: (2026)
von: Yang, Tao, et al.
Veröffentlicht: (2026)
GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
von: Zou, Hao, et al.
Veröffentlicht: (2025)
von: Zou, Hao, et al.
Veröffentlicht: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
von: Lee, Jihoon, et al.
Veröffentlicht: (2025)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2026)
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2026)
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
von: Liu, Yating, et al.
Veröffentlicht: (2023)
von: Liu, Yating, et al.
Veröffentlicht: (2023)
DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding
von: Xu, Jingqi
Veröffentlicht: (2026)
von: Xu, Jingqi
Veröffentlicht: (2026)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
von: Gui, Rui, et al.
Veröffentlicht: (2025)
von: Gui, Rui, et al.
Veröffentlicht: (2025)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
von: Kim, Kwonyoung, et al.
Veröffentlicht: (2025)
von: Kim, Kwonyoung, et al.
Veröffentlicht: (2025)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
Separate-and-Enhance: Compositional Finetuning for Text2Image Diffusion Models
von: Bao, Zhipeng, et al.
Veröffentlicht: (2023)
von: Bao, Zhipeng, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
von: Kwon, Jihoon, et al.
Veröffentlicht: (2026) -
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
von: Jeon, Jaebyeong, et al.
Veröffentlicht: (2025) -
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
von: Shivika, et al.
Veröffentlicht: (2026) -
IDEA: Image Description Enhanced CLIP-Adapter
von: Ye, Zhipeng, et al.
Veröffentlicht: (2025) -
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2024)