Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yafei, Shang, Yongle, Li, Huafeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
by: Ouyang, Pengxiang, et al.
Published: (2025)
by: Ouyang, Pengxiang, et al.
Published: (2025)
Cross-Modal Coordination Across a Diverse Set of Input Modalities
by: Sánchez, Jorge, et al.
Published: (2024)
by: Sánchez, Jorge, et al.
Published: (2024)
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
by: Zhang, Shiyi, et al.
Published: (2026)
by: Zhang, Shiyi, et al.
Published: (2026)
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking
by: Ahtesham, Muhammad, et al.
Published: (2025)
by: Ahtesham, Muhammad, et al.
Published: (2025)
Deep Learning-based Text-in-Image Watermarking
by: Karki, Bishwa, et al.
Published: (2024)
by: Karki, Bishwa, et al.
Published: (2024)
Knowledge Bridger: Towards Training-free Missing Modality Completion
by: Ke, Guanzhou, et al.
Published: (2025)
by: Ke, Guanzhou, et al.
Published: (2025)
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
Overcome Modal Bias in Multi-modal Federated Learning via Balanced Modality Selection
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Operationalizing Fairness in Text-to-Image Models: A Survey of Bias, Fairness Audits and Mitigation Strategies
by: Smith, Megan, et al.
Published: (2026)
by: Smith, Megan, et al.
Published: (2026)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
by: Saleh, Mohamed, et al.
Published: (2026)
by: Saleh, Mohamed, et al.
Published: (2026)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
by: Zhang, Zheng, et al.
Published: (2024)
by: Zhang, Zheng, et al.
Published: (2024)
MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates
by: Zhao, Binyu, et al.
Published: (2025)
by: Zhao, Binyu, et al.
Published: (2025)
Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
by: Liu, Jiale, et al.
Published: (2025)
by: Liu, Jiale, et al.
Published: (2025)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention
by: Panta, Sanjeev, et al.
Published: (2026)
by: Panta, Sanjeev, et al.
Published: (2026)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
Long-Range Feature Propagating for Natural Image Matting
by: Liu, Qinglin, et al.
Published: (2021)
by: Liu, Qinglin, et al.
Published: (2021)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding
by: Madan, Surbhi, et al.
Published: (2024)
by: Madan, Surbhi, et al.
Published: (2024)
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation
by: Nag, Sayan, et al.
Published: (2024)
by: Nag, Sayan, et al.
Published: (2024)
HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
by: Gao, Lishuai, et al.
Published: (2024)
by: Gao, Lishuai, et al.
Published: (2024)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
by: Shang, Ziqiao, et al.
Published: (2023)
by: Shang, Ziqiao, et al.
Published: (2023)
GroMo: Plant Growth Modeling with Multiview Images
by: Bhatt, Ruchi, et al.
Published: (2025)
by: Bhatt, Ruchi, et al.
Published: (2025)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
by: Yang, Shuyu, et al.
Published: (2024)
by: Yang, Shuyu, et al.
Published: (2024)
Bridging Compressed Image Latents and Multimodal Large Language Models
by: Kao, Chia-Hao, et al.
Published: (2024)
by: Kao, Chia-Hao, et al.
Published: (2024)
Residual Prior-driven Frequency-aware Network for Image Fusion
by: Zheng, Guan, et al.
Published: (2025)
by: Zheng, Guan, et al.
Published: (2025)
An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture
by: Shi, Tiancheng, et al.
Published: (2024)
by: Shi, Tiancheng, et al.
Published: (2024)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Similar Items
-
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
by: Huang, Hailang, et al.
Published: (2024) -
PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
by: Ouyang, Pengxiang, et al.
Published: (2025) -
Cross-Modal Coordination Across a Diverse Set of Input Modalities
by: Sánchez, Jorge, et al.
Published: (2024) -
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
by: Li, Hui, et al.
Published: (2025) -
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)