Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Le, Awal, Rabiul, Agrawal, Aishwarya |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Benchmarking Vision Language Models for Cultural Understanding
by: Nayak, Shravan, et al.
Published: (2024)
by: Nayak, Shravan, et al.
Published: (2024)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
by: Zhang, Jianrui, et al.
Published: (2024)
by: Zhang, Jianrui, et al.
Published: (2024)
Dual-Level Cross-Modal Contrastive Clustering
by: Zhang, Haixin, et al.
Published: (2024)
by: Zhang, Haixin, et al.
Published: (2024)
IntraStyler: Intra-Domain Style Synthesis for Cross-Modality MRI Domain Adaptation
by: Liu, Han, et al.
Published: (2026)
by: Liu, Han, et al.
Published: (2026)
Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies
by: Tian, Mulin, et al.
Published: (2023)
by: Tian, Mulin, et al.
Published: (2023)
Enhancing Cross-Modal Medical Image Segmentation through Compositionality
by: Eijpe, Aniek, et al.
Published: (2024)
by: Eijpe, Aniek, et al.
Published: (2024)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
Assessing and Learning Alignment of Unimodal Vision and Language Models
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
by: Zhang, Le, et al.
Published: (2026)
by: Zhang, Le, et al.
Published: (2026)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
by: Herzog, Jonas, et al.
Published: (2026)
by: Herzog, Jonas, et al.
Published: (2026)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
by: Jiang, Hanqi, et al.
Published: (2024)
by: Jiang, Hanqi, et al.
Published: (2024)
Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
by: Kim, SiWoo, et al.
Published: (2025)
by: Kim, SiWoo, et al.
Published: (2025)
PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities
by: Chen, Jiajun, et al.
Published: (2025)
by: Chen, Jiajun, et al.
Published: (2025)
Enhancing Contrastive Learning for Geolocalization by Discovering Hard Negatives on Semivariograms
by: Chen, Boyi, et al.
Published: (2025)
by: Chen, Boyi, et al.
Published: (2025)
Breaking the Geometric Bottleneck: Contrastive Expansion in Asymmetric Cross-Modal Distillation
by: Thayani, Kabir
Published: (2026)
by: Thayani, Kabir
Published: (2026)
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
by: Xue, Youze, et al.
Published: (2025)
by: Xue, Youze, et al.
Published: (2025)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
Contrast-X: A Multi-Modal Contrast Image Synthesis Benchmark and Universal Modality Flow Matching
by: Chen, Yifan, et al.
Published: (2026)
by: Chen, Yifan, et al.
Published: (2026)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
by: Ma, Wenxuan, et al.
Published: (2024)
by: Ma, Wenxuan, et al.
Published: (2024)
Understanding Hyperbolic Metric Learning through Hard Negative Sampling
by: Yue, Yun, et al.
Published: (2024)
by: Yue, Yun, et al.
Published: (2024)
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
by: Zhu, Tianyu, et al.
Published: (2024)
by: Zhu, Tianyu, et al.
Published: (2024)
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
by: Liu, Zheng, et al.
Published: (2025)
by: Liu, Zheng, et al.
Published: (2025)
Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation
by: Cai, Lincan, et al.
Published: (2024)
by: Cai, Lincan, et al.
Published: (2024)
Manipulating Multimodal Agents via Cross-Modal Prompt Injection
by: Wang, Le, et al.
Published: (2025)
by: Wang, Le, et al.
Published: (2025)
Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
by: Jiang, Tianjiao, et al.
Published: (2025)
by: Jiang, Tianjiao, et al.
Published: (2025)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
by: Trinh, Quoc-Huy, et al.
Published: (2026)
by: Trinh, Quoc-Huy, et al.
Published: (2026)
CPCL: Cross-Modal Prototypical Contrastive Learning for Weakly Supervised Text-based Person Retrieval
by: Zhao, Xinpeng, et al.
Published: (2024)
by: Zhao, Xinpeng, et al.
Published: (2024)
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
by: Shi, Yi, et al.
Published: (2025)
by: Shi, Yi, et al.
Published: (2025)
Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition
by: Radevski, Gorjan
Published: (2025)
by: Radevski, Gorjan
Published: (2025)
Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples
by: Rösch, Philipp J., et al.
Published: (2024)
by: Rösch, Philipp J., et al.
Published: (2024)
A Generalization Theory of Cross-Modality Distillation with Contrastive Learning
by: Lin, Hangyu, et al.
Published: (2024)
by: Lin, Hangyu, et al.
Published: (2024)
Locality-Sensitive Hashing for Efficient Hard Negative Sampling in Contrastive Learning
by: Deuser, Fabian, et al.
Published: (2025)
by: Deuser, Fabian, et al.
Published: (2025)
Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
by: Li, Zihan, et al.
Published: (2025)
by: Li, Zihan, et al.
Published: (2025)
Similar Items
-
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023) -
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024) -
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
by: Oh, Youngtaek, et al.
Published: (2024) -
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025) -
Benchmarking Vision Language Models for Cultural Understanding
by: Nayak, Shravan, et al.
Published: (2024)