Descriptive Image-Text Matching with Graded Contextual Similarity
Fuente:
arXiv
Saved in:
| Main Authors: | Jang, Jinhyun, Lee, Jiyoung, Sohn, Kwanghoon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Visual Recognition with Hyperbolical Visual Hierarchy Mapping
by: Kwon, Hyeongjun, et al.
Published: (2024)
by: Kwon, Hyeongjun, et al.
Published: (2024)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
Bridging Vision and Language Spaces with Assignment Prediction
by: Park, Jungin, et al.
Published: (2024)
by: Park, Jungin, et al.
Published: (2024)
Saliency-Aware Model Merging
by: Park, Jungin, et al.
Published: (2026)
by: Park, Jungin, et al.
Published: (2026)
V-LynX: Token Interface Alignment for Video+X LLMs
by: Park, Jungin, et al.
Published: (2026)
by: Park, Jungin, et al.
Published: (2026)
PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation
by: Kim, Kwonyoung, et al.
Published: (2022)
by: Kim, Kwonyoung, et al.
Published: (2022)
EBDM: Exemplar-guided Image Translation with Brownian-bridge Diffusion Models
by: Lee, Eungbean, et al.
Published: (2024)
by: Lee, Eungbean, et al.
Published: (2024)
Diffusion-driven GAN Inversion for Multi-Modal Face Image Generation
by: Kim, Jihyun, et al.
Published: (2024)
by: Kim, Jihyun, et al.
Published: (2024)
Rethinking Open-World Semi-Supervised Learning: Distribution Mismatch and Inductive Inference
by: Park, Seongheon, et al.
Published: (2024)
by: Park, Seongheon, et al.
Published: (2024)
Enhancing Source-Free Domain Adaptive Object Detection with Low-confidence Pseudo Label Distillation
by: Yoon, Ilhoon, et al.
Published: (2024)
by: Yoon, Ilhoon, et al.
Published: (2024)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
by: Kim, Kwonyoung, et al.
Published: (2025)
by: Kim, Kwonyoung, et al.
Published: (2025)
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
by: Kwon, Jihoon, et al.
Published: (2025)
by: Kwon, Jihoon, et al.
Published: (2025)
DreamFlow: High-Quality Text-to-3D Generation by Approximating Probability Flow
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
by: Lim, Jiyoung, et al.
Published: (2026)
by: Lim, Jiyoung, et al.
Published: (2026)
ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression
by: Lee, Juneyong, et al.
Published: (2026)
by: Lee, Juneyong, et al.
Published: (2026)
Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild
by: Jin, Siyoon, et al.
Published: (2024)
by: Jin, Siyoon, et al.
Published: (2024)
DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization
by: Jang, Geonhui, et al.
Published: (2024)
by: Jang, Geonhui, et al.
Published: (2024)
HySim: An Efficient Hybrid Similarity Measure for Patch Matching in Image Inpainting
by: Noufel, Saad, et al.
Published: (2024)
by: Noufel, Saad, et al.
Published: (2024)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
by: Seo, Junyoung, et al.
Published: (2023)
by: Seo, Junyoung, et al.
Published: (2023)
CountSteer: Steering Attention for Object Counting in Diffusion Models
by: Boo, Hyemin, et al.
Published: (2025)
by: Boo, Hyemin, et al.
Published: (2025)
MASS: Overcoming Language Bias in Image-Text Matching
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation
by: Wei, Tianyi, et al.
Published: (2024)
by: Wei, Tianyi, et al.
Published: (2024)
Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects
by: Qiu, Weimin, et al.
Published: (2024)
by: Qiu, Weimin, et al.
Published: (2024)
Compositional Image-Text Matching and Retrieval by Grounding Entities
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization
by: Nam, Jisu, et al.
Published: (2024)
by: Nam, Jisu, et al.
Published: (2024)
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)
by: Lee, Junho, et al.
Published: (2026)
Matching Semantically Similar Non-Identical Objects
by: Marumo, Yusuke, et al.
Published: (2024)
by: Marumo, Yusuke, et al.
Published: (2024)
Exploiting Text-Image Latent Spaces for the Description of Visual Concepts
by: Schmalwasser, Laines, et al.
Published: (2024)
by: Schmalwasser, Laines, et al.
Published: (2024)
Referee: Reference-aware Audiovisual Deepfake Detection
by: Boo, Hyemin, et al.
Published: (2025)
by: Boo, Hyemin, et al.
Published: (2025)
CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval
by: Kang, Bin, et al.
Published: (2025)
by: Kang, Bin, et al.
Published: (2025)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
by: Yue, Xinli, et al.
Published: (2025)
by: Yue, Xinli, et al.
Published: (2025)
Affective Image Editing: Shaping Emotional Factors via Text Descriptions
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
by: Ma, Leilei, et al.
Published: (2024)
by: Ma, Leilei, et al.
Published: (2024)
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
by: Lin, Honglin, et al.
Published: (2024)
by: Lin, Honglin, et al.
Published: (2024)
Identity Decoupling for Multi-Subject Personalization of Text-to-Image Models
by: Jang, Sangwon, et al.
Published: (2024)
by: Jang, Sangwon, et al.
Published: (2024)
Direct Unlearning Optimization for Robust and Safe Text-to-Image Models
by: Park, Yong-Hyun, et al.
Published: (2024)
by: Park, Yong-Hyun, et al.
Published: (2024)
Similar Items
-
Improving Visual Recognition with Hyperbolical Visual Hierarchy Mapping
by: Kwon, Hyeongjun, et al.
Published: (2024) -
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025) -
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025) -
Bridging Vision and Language Spaces with Assignment Prediction
by: Park, Jungin, et al.
Published: (2024) -
Saliency-Aware Model Merging
by: Park, Jungin, et al.
Published: (2026)