Compositional Image Retrieval via Instruction-Aware Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Wenliang, An, Weizhi, Jiang, Feng, Ma, Hehuan, Guo, Yuzhi, Huang, Junzhou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning
by: Zhou, Qifeng, et al.
Published: (2024)
by: Zhou, Qifeng, et al.
Published: (2024)
HOMIE: Histopathology Omni-modal Embedding for Pathology Composed Retrieval
by: Zhou, Qifeng, et al.
Published: (2025)
by: Zhou, Qifeng, et al.
Published: (2025)
Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation
by: Na, Saiyang, et al.
Published: (2024)
by: Na, Saiyang, et al.
Published: (2024)
Leveraging Gait Patterns as Biomarkers: An attention-guided Deep Multiple Instance Learning Network for Scoliosis Classification
by: Li, Haiqing, et al.
Published: (2025)
by: Li, Haiqing, et al.
Published: (2025)
Text-Guided Multi-Instance Learning for Scoliosis Screening via Gait Video Analysis
by: Li, Haiqing, et al.
Published: (2025)
by: Li, Haiqing, et al.
Published: (2025)
DELST: Dual Entailment Learning for Hyperbolic Image-Gene Pretraining in Spatial Transcriptomics
by: Chen, Xulin, et al.
Published: (2025)
by: Chen, Xulin, et al.
Published: (2025)
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
by: Guo, Wenliang, et al.
Published: (2025)
by: Guo, Wenliang, et al.
Published: (2025)
Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives
by: Feng, Zhangchi, et al.
Published: (2024)
by: Feng, Zhangchi, et al.
Published: (2024)
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
by: Zhong, Wenliang, et al.
Published: (2024)
by: Zhong, Wenliang, et al.
Published: (2024)
CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
by: Wang, Haoran, et al.
Published: (2022)
by: Wang, Haoran, et al.
Published: (2022)
Adaptive Noise-Tolerant Network for Image Segmentation
by: Li, Weizhi
Published: (2025)
by: Li, Weizhi
Published: (2025)
Lightweight Contrastive Distilled Hashing for Online Cross-modal Retrieval
by: Li, Jiaxing, et al.
Published: (2025)
by: Li, Jiaxing, et al.
Published: (2025)
ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling
by: Mao, Chaojie, et al.
Published: (2025)
by: Mao, Chaojie, et al.
Published: (2025)
Transformer-based Clipped Contrastive Quantization Learning for Unsupervised Image Retrieval
by: Dubey, Ayush, et al.
Published: (2024)
by: Dubey, Ayush, et al.
Published: (2024)
NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval
by: Gao, Peng, et al.
Published: (2025)
by: Gao, Peng, et al.
Published: (2025)
Procedural Mistake Detection via Action Effect Modeling
by: Guo, Wenliang, et al.
Published: (2025)
by: Guo, Wenliang, et al.
Published: (2025)
C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning
by: Ma, Ji, et al.
Published: (2024)
by: Ma, Ji, et al.
Published: (2024)
VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
by: Ma, Jian, et al.
Published: (2025)
by: Ma, Jian, et al.
Published: (2025)
LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer
by: Chen, Yuzhuo, et al.
Published: (2025)
by: Chen, Yuzhuo, et al.
Published: (2025)
LeOCLR: Leveraging Original Images for Contrastive Learning of Visual Representations
by: Alkhalefi, Mohammad, et al.
Published: (2024)
by: Alkhalefi, Mohammad, et al.
Published: (2024)
Style-Preserving Lip Sync via Audio-Aware Style Reference
by: Zhong, Weizhi, et al.
Published: (2024)
by: Zhong, Weizhi, et al.
Published: (2024)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Learning from History: Task-agnostic Model Contrastive Learning for Image Restoration
by: Wu, Gang, et al.
Published: (2023)
by: Wu, Gang, et al.
Published: (2023)
Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
by: Jin, Jiayun, et al.
Published: (2026)
by: Jin, Jiayun, et al.
Published: (2026)
Learning to Restore Multi-Degraded Images via Ingredient Decoupling and Task-Aware Path Adaptation
by: Gao, Hu, et al.
Published: (2025)
by: Gao, Hu, et al.
Published: (2025)
Underwater Image Enhancement with Cascaded Contrastive Learning
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models
by: Qiao, Xiaozhen, et al.
Published: (2025)
by: Qiao, Xiaozhen, et al.
Published: (2025)
Video-based Generalized Category Discovery via Memory-Guided Consistency-Aware Contrastive Learning
by: Jing, Zhang, et al.
Published: (2025)
by: Jing, Zhang, et al.
Published: (2025)
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning
by: Chen, Sherry X., et al.
Published: (2025)
by: Chen, Sherry X., et al.
Published: (2025)
FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
Learn From Zoom: Decoupled Supervised Contrastive Learning For WCE Image Classification
by: Qiu, Kunpeng, et al.
Published: (2024)
by: Qiu, Kunpeng, et al.
Published: (2024)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
Dynamic Contrastive Learning for Hierarchical Retrieval: A Case Study of Distance-Aware Cross-View Geo-Localization
by: Zhang, Suofei, et al.
Published: (2025)
by: Zhang, Suofei, et al.
Published: (2025)
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
by: Zeng, Yingsen, et al.
Published: (2024)
by: Zeng, Yingsen, et al.
Published: (2024)
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
by: Wang, Chenglin, et al.
Published: (2025)
by: Wang, Chenglin, et al.
Published: (2025)
Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image Restoration
by: Wu, Gang, et al.
Published: (2025)
by: Wu, Gang, et al.
Published: (2025)
Similar Items
-
PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning
by: Zhou, Qifeng, et al.
Published: (2024) -
HOMIE: Histopathology Omni-modal Embedding for Pathology Composed Retrieval
by: Zhou, Qifeng, et al.
Published: (2025) -
Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation
by: Na, Saiyang, et al.
Published: (2024) -
Leveraging Gait Patterns as Biomarkers: An attention-guided Deep Multiple Instance Learning Network for Scoliosis Classification
by: Li, Haiqing, et al.
Published: (2025) -
Text-Guided Multi-Instance Learning for Scoliosis Screening via Gait Video Analysis
by: Li, Haiqing, et al.
Published: (2025)