PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Ouyang, Pengxiang, Ma, Qing, Wang, Zheng, Bai, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
by: Zhang, Yafei, et al.
Published: (2025)
by: Zhang, Yafei, et al.
Published: (2025)
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
by: Ning, Hailong, et al.
Published: (2025)
by: Ning, Hailong, et al.
Published: (2025)
Residual Prior-driven Frequency-aware Network for Image Fusion
by: Zheng, Guan, et al.
Published: (2025)
by: Zheng, Guan, et al.
Published: (2025)
Semantic-Aware Adversarial Training for Reliable Deep Hashing Retrieval
by: Yuan, Xu, et al.
Published: (2023)
by: Yuan, Xu, et al.
Published: (2023)
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking
by: Ahtesham, Muhammad, et al.
Published: (2025)
by: Ahtesham, Muhammad, et al.
Published: (2025)
Deep Learning-based Text-in-Image Watermarking
by: Karki, Bishwa, et al.
Published: (2024)
by: Karki, Bishwa, et al.
Published: (2024)
Operationalizing Fairness in Text-to-Image Models: A Survey of Bias, Fairness Audits and Mitigation Strategies
by: Smith, Megan, et al.
Published: (2026)
by: Smith, Megan, et al.
Published: (2026)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework
by: Karaca, Ali Can, et al.
Published: (2025)
by: Karaca, Ali Can, et al.
Published: (2025)
Pseudo-triplet Guided Few-shot Composed Image Retrieval
by: Hou, Bohan, et al.
Published: (2024)
by: Hou, Bohan, et al.
Published: (2024)
ClassWise-CRF: Category-Specific Fusion for Enhanced Semantic Segmentation of Remote Sensing Imagery
by: Zhu, Qinfeng, et al.
Published: (2025)
by: Zhu, Qinfeng, et al.
Published: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
by: Gao, Lishuai, et al.
Published: (2024)
by: Gao, Lishuai, et al.
Published: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
by: Duan, Yue, et al.
Published: (2024)
by: Duan, Yue, et al.
Published: (2024)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
by: Ye, Chengyang, et al.
Published: (2024)
by: Ye, Chengyang, et al.
Published: (2024)
MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic Facial Expression Recognition
by: Xiang, Peihao, et al.
Published: (2024)
by: Xiang, Peihao, et al.
Published: (2024)
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
by: Pegia, Maria-Eirini, et al.
Published: (2026)
by: Pegia, Maria-Eirini, et al.
Published: (2026)
Copy-Move Forgery Detection and Question Answering for Remote Sensing Image
by: Zhang, Ze, et al.
Published: (2024)
by: Zhang, Ze, et al.
Published: (2024)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
MB-ORES: A Multi-Branch Object Reasoner for Visual Grounding in Remote Sensing
by: Radouane, Karim, et al.
Published: (2025)
by: Radouane, Karim, et al.
Published: (2025)
360VFI: A Dataset and Benchmark for Omnidirectional Video Frame Interpolation
by: Lu, Wenxuan, et al.
Published: (2024)
by: Lu, Wenxuan, et al.
Published: (2024)
CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos
by: Panambur, Tejas, et al.
Published: (2025)
by: Panambur, Tejas, et al.
Published: (2025)
Long-Range Feature Propagating for Natural Image Matting
by: Liu, Qinglin, et al.
Published: (2021)
by: Liu, Qinglin, et al.
Published: (2021)
GroMo: Plant Growth Modeling with Multiview Images
by: Bhatt, Ruchi, et al.
Published: (2025)
by: Bhatt, Ruchi, et al.
Published: (2025)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture
by: Shi, Tiancheng, et al.
Published: (2024)
by: Shi, Tiancheng, et al.
Published: (2024)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval
by: Moummad, Ilyass, et al.
Published: (2026)
by: Moummad, Ilyass, et al.
Published: (2026)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
Flow Generator Matching
by: Huang, Zemin, et al.
Published: (2024)
by: Huang, Zemin, et al.
Published: (2024)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
by: Wang, Ziteng, et al.
Published: (2026)
by: Wang, Ziteng, et al.
Published: (2026)
InvZW: Invariant Feature Learning via Noise-Adversarial Training for Robust Image Zero-Watermarking
by: Tanvir, Abdullah All, et al.
Published: (2025)
by: Tanvir, Abdullah All, et al.
Published: (2025)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Similar Items
-
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
by: Zhang, Yafei, et al.
Published: (2025) -
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
by: Yang, Rui, et al.
Published: (2024) -
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024) -
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
by: Ning, Hailong, et al.
Published: (2025) -
Residual Prior-driven Frequency-aware Network for Image Fusion
by: Zheng, Guan, et al.
Published: (2025)