Generalized Contrastive Learning for Universal Multimodal Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jungsoo, Cho, Janghoon, Park, Hyojin, Hayat, Munawar, Hwang, Kyuwoong, Porikli, Fatih, Choi, Sungha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025)
by: Cho, Janghoon, et al.
Published: (2025)
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
by: Lee, Jungsoo, et al.
Published: (2025)
by: Lee, Jungsoo, et al.
Published: (2025)
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
by: Park, Sunghyun, et al.
Published: (2025)
by: Park, Sunghyun, et al.
Published: (2025)
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
by: Park, Hyojin, et al.
Published: (2026)
by: Park, Hyojin, et al.
Published: (2026)
Attention Guided Alignment in Efficient Vision-Language Models
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
by: Garrepalli, Risheek, et al.
Published: (2024)
by: Garrepalli, Risheek, et al.
Published: (2024)
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
Resolving the Identity Crisis in Text-to-Image Generation
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
PosSAM: Panoptic Open-vocabulary Segment Anything
by: VS, Vibashan, et al.
Published: (2024)
by: VS, Vibashan, et al.
Published: (2024)
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
by: Kadambi, Shreya, et al.
Published: (2025)
by: Kadambi, Shreya, et al.
Published: (2025)
Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning
by: Lee, Juntae, et al.
Published: (2025)
by: Lee, Juntae, et al.
Published: (2025)
ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
by: Das, Debasmit, et al.
Published: (2025)
by: Das, Debasmit, et al.
Published: (2025)
Feature Diversification and Adaptation for Federated Domain Generalization
by: Yang, Seunghan, et al.
Published: (2024)
by: Yang, Seunghan, et al.
Published: (2024)
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
Flashbacks to Harmonize Stability and Plasticity in Continual Learning
by: Mahmoodi, Leila, et al.
Published: (2025)
by: Mahmoodi, Leila, et al.
Published: (2025)
Performance Plateaus in Inference-Time Scaling for Text-to-Image Diffusion Without External Models
by: Choi, Changhyun, et al.
Published: (2025)
by: Choi, Changhyun, et al.
Published: (2025)
OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware Interpolation
by: Jeong, Jisoo, et al.
Published: (2024)
by: Jeong, Jisoo, et al.
Published: (2024)
Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models
by: Cho, Wonguk, et al.
Published: (2024)
by: Cho, Wonguk, et al.
Published: (2024)
SMCL: Saliency Masked Contrastive Learning for Long-tailed Recognition
by: Park, Sanglee, et al.
Published: (2024)
by: Park, Sanglee, et al.
Published: (2024)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
by: Park, Sunghyun, et al.
Published: (2026)
by: Park, Sunghyun, et al.
Published: (2026)
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
by: Porikli, Sedat, et al.
Published: (2025)
by: Porikli, Sedat, et al.
Published: (2025)
Collaborative Learning with Multiple Foundation Models for Source-Free Domain Adaptation
by: Lee, Huisoo, et al.
Published: (2025)
by: Lee, Huisoo, et al.
Published: (2025)
Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning
by: Du, Ruoshuang, et al.
Published: (2026)
by: Du, Ruoshuang, et al.
Published: (2026)
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
by: Zhu, Tianyu, et al.
Published: (2024)
by: Zhu, Tianyu, et al.
Published: (2024)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
by: Bahng, Hyojin, et al.
Published: (2025)
by: Bahng, Hyojin, et al.
Published: (2025)
Multimodal Transformer With a Low-Computational-Cost Guarantee
by: Park, Sungjin, et al.
Published: (2024)
by: Park, Sungjin, et al.
Published: (2024)
Leveraging Programmatically Generated Synthetic Data for Differentially Private Diffusion Training
by: Choi, Yujin, et al.
Published: (2024)
by: Choi, Yujin, et al.
Published: (2024)
Anomaly Score: Evaluating Generative Models and Individual Generated Images based on Complexity and Vulnerability
by: Hwang, Jaehui, et al.
Published: (2023)
by: Hwang, Jaehui, et al.
Published: (2023)
Differential-informed Sample Selection Accelerates Multimodal Contrastive Learning
by: Zhao, Zihua, et al.
Published: (2025)
by: Zhao, Zihua, et al.
Published: (2025)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
by: Park, Yeji, et al.
Published: (2024)
by: Park, Yeji, et al.
Published: (2024)
$\mathbb{X}$-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs
by: Sobal, Vlad, et al.
Published: (2024)
by: Sobal, Vlad, et al.
Published: (2024)
Multimodal Unsupervised Domain Generalization by Retrieving Across the Modality Gap
by: Liao, Christopher, et al.
Published: (2024)
by: Liao, Christopher, et al.
Published: (2024)
CLIPLoss and Norm-Based Data Selection Methods for Multimodal Contrastive Learning
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data
by: Jing, Long, et al.
Published: (2026)
by: Jing, Long, et al.
Published: (2026)
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
by: Jo, Dae Ung, et al.
Published: (2024)
by: Jo, Dae Ung, et al.
Published: (2024)
Similar Items
-
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025) -
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
by: Lee, Jungsoo, et al.
Published: (2025) -
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
by: Park, Sunghyun, et al.
Published: (2025) -
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
by: Park, Minho, et al.
Published: (2025) -
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
by: Park, Hyojin, et al.
Published: (2026)