Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
Fuente:
arXiv
Guardado en:
| Autores principales: | Kwon, Jihoon, Min, Kyle, Sohn, Jy-yong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
por: Kwon, Jihoon, et al.
Publicado: (2026)
por: Kwon, Jihoon, et al.
Publicado: (2026)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
por: Jeon, Jaebyeong, et al.
Publicado: (2025)
por: Jeon, Jaebyeong, et al.
Publicado: (2025)
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
por: Shivika, et al.
Publicado: (2026)
por: Shivika, et al.
Publicado: (2026)
IDEA: Image Description Enhanced CLIP-Adapter
por: Ye, Zhipeng, et al.
Publicado: (2025)
por: Ye, Zhipeng, et al.
Publicado: (2025)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
por: Sbrolli, Cristian, et al.
Publicado: (2024)
por: Sbrolli, Cristian, et al.
Publicado: (2024)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
por: Che, Chang, et al.
Publicado: (2024)
por: Che, Chang, et al.
Publicado: (2024)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
por: Asokan, Mothilal, et al.
Publicado: (2025)
por: Asokan, Mothilal, et al.
Publicado: (2025)
ComCLIP: Training-Free Compositional Image and Text Matching
por: Jiang, Kenan, et al.
Publicado: (2022)
por: Jiang, Kenan, et al.
Publicado: (2022)
A Theoretical Framework for Preventing Class Collapse in Supervised Contrastive Learning
por: Lee, Chungpa, et al.
Publicado: (2025)
por: Lee, Chungpa, et al.
Publicado: (2025)
Interpreting CLIP's Image Representation via Text-Based Decomposition
por: Gandelsman, Yossi, et al.
Publicado: (2023)
por: Gandelsman, Yossi, et al.
Publicado: (2023)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
por: Cao, Anh-Quan, et al.
Publicado: (2024)
por: Cao, Anh-Quan, et al.
Publicado: (2024)
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
por: Ma, Wenxin, et al.
Publicado: (2025)
por: Ma, Wenxin, et al.
Publicado: (2025)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
por: Prakash, Nirmalendu, et al.
Publicado: (2026)
por: Prakash, Nirmalendu, et al.
Publicado: (2026)
Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis
por: Samanta, Argha Kamal, et al.
Publicado: (2025)
por: Samanta, Argha Kamal, et al.
Publicado: (2025)
FG-CLIP: Fine-Grained Visual and Textual Alignment
por: Xie, Chunyu, et al.
Publicado: (2025)
por: Xie, Chunyu, et al.
Publicado: (2025)
Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
por: Cao, Min, et al.
Publicado: (2025)
por: Cao, Min, et al.
Publicado: (2025)
Parrot Captions Teach CLIP to Spot Text
por: Lin, Yiqi, et al.
Publicado: (2023)
por: Lin, Yiqi, et al.
Publicado: (2023)
Updating CLIP to Prefer Descriptions Over Captions
por: Zur, Amir, et al.
Publicado: (2024)
por: Zur, Amir, et al.
Publicado: (2024)
V-LynX: Token Interface Alignment for Video+X LLMs
por: Park, Jungin, et al.
Publicado: (2026)
por: Park, Jungin, et al.
Publicado: (2026)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
por: He, Chiyuan, et al.
Publicado: (2025)
por: He, Chiyuan, et al.
Publicado: (2025)
DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment
por: Chen, Weizhi, et al.
Publicado: (2025)
por: Chen, Weizhi, et al.
Publicado: (2025)
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
por: Xie, Shaoan, et al.
Publicado: (2025)
por: Xie, Shaoan, et al.
Publicado: (2025)
Enhancing Source-Free Domain Adaptive Object Detection with Low-confidence Pseudo Label Distillation
por: Yoon, Ilhoon, et al.
Publicado: (2024)
por: Yoon, Ilhoon, et al.
Publicado: (2024)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
por: Kim, Jiyeong, et al.
Publicado: (2026)
por: Kim, Jiyeong, et al.
Publicado: (2026)
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
por: Valdez, Hector A., et al.
Publicado: (2024)
por: Valdez, Hector A., et al.
Publicado: (2024)
Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
por: Xue, Leyan, et al.
Publicado: (2025)
por: Xue, Leyan, et al.
Publicado: (2025)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
por: Yang, Shengzhu, et al.
Publicado: (2025)
por: Yang, Shengzhu, et al.
Publicado: (2025)
Discriminative Perception via Anchored Description for Reasoning Segmentation
por: Yang, Tao, et al.
Publicado: (2026)
por: Yang, Tao, et al.
Publicado: (2026)
GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
por: Zou, Hao, et al.
Publicado: (2025)
por: Zou, Hao, et al.
Publicado: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
por: Lee, Jihoon, et al.
Publicado: (2025)
por: Lee, Jihoon, et al.
Publicado: (2025)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
por: Lee, Jaa-Yeon, et al.
Publicado: (2026)
por: Lee, Jaa-Yeon, et al.
Publicado: (2026)
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
por: Vatsa, Mayank, et al.
Publicado: (2025)
por: Vatsa, Mayank, et al.
Publicado: (2025)
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
por: Liu, Yating, et al.
Publicado: (2023)
por: Liu, Yating, et al.
Publicado: (2023)
DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
por: Liu, Fuxiao, et al.
Publicado: (2023)
por: Liu, Fuxiao, et al.
Publicado: (2023)
Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding
por: Xu, Jingqi
Publicado: (2026)
por: Xu, Jingqi
Publicado: (2026)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
por: Xu, Wenhao, et al.
Publicado: (2024)
por: Xu, Wenhao, et al.
Publicado: (2024)
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
por: Gui, Rui, et al.
Publicado: (2025)
por: Gui, Rui, et al.
Publicado: (2025)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
por: Kim, Kwonyoung, et al.
Publicado: (2025)
por: Kim, Kwonyoung, et al.
Publicado: (2025)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
por: Zeng, Gangyan, et al.
Publicado: (2024)
por: Zeng, Gangyan, et al.
Publicado: (2024)
Separate-and-Enhance: Compositional Finetuning for Text2Image Diffusion Models
por: Bao, Zhipeng, et al.
Publicado: (2023)
por: Bao, Zhipeng, et al.
Publicado: (2023)
Ejemplares similares
-
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
por: Kwon, Jihoon, et al.
Publicado: (2026) -
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
por: Jeon, Jaebyeong, et al.
Publicado: (2025) -
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
por: Shivika, et al.
Publicado: (2026) -
IDEA: Image Description Enhanced CLIP-Adapter
por: Ye, Zhipeng, et al.
Publicado: (2025) -
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
por: Sbrolli, Cristian, et al.
Publicado: (2024)