Improving Text-based Person Search via Part-level Cross-modal Correspondence
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jicheol, Jeong, Boseung, Kim, Dongwon, Kwak, Suha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
by: Park, Jicheol, et al.
Published: (2024)
by: Park, Jicheol, et al.
Published: (2024)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
by: Kim, Sungyeon, et al.
Published: (2024)
by: Kim, Sungyeon, et al.
Published: (2024)
Bootstrapping Top-down Information for Self-modulating Slot Attention
by: Kim, Dongwon, et al.
Published: (2024)
by: Kim, Dongwon, et al.
Published: (2024)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)
by: Kim, Inho, et al.
Published: (2025)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
by: Kim, Seoyeon, et al.
Published: (2023)
by: Kim, Seoyeon, et al.
Published: (2023)
Improving Robustness to Multiple Spurious Correlations by Multi-Objective Optimization
by: Kim, Nayeong, et al.
Published: (2024)
by: Kim, Nayeong, et al.
Published: (2024)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
by: Kim, Sungyeon, et al.
Published: (2023)
by: Kim, Sungyeon, et al.
Published: (2023)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
by: Gong, Dayoung, et al.
Published: (2024)
by: Gong, Dayoung, et al.
Published: (2024)
Cross-modal Active Complementary Learning with Self-refining Correspondence
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
Enhancing Cost Efficiency in Active Learning with Candidate Set Query
by: Gwon, Yeho, et al.
Published: (2025)
by: Gwon, Yeho, et al.
Published: (2025)
Memory-Efficient Personalization of Text-to-Image Diffusion Models via Selective Optimization Strategies
by: Choi, Seokeon, et al.
Published: (2025)
by: Choi, Seokeon, et al.
Published: (2025)
GENIUS: A Generative Framework for Universal Multimodal Search
by: Kim, Sungyeon, et al.
Published: (2025)
by: Kim, Sungyeon, et al.
Published: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
by: Kim, Dongkeun, et al.
Published: (2025)
by: Kim, Dongkeun, et al.
Published: (2025)
Directional Textual Inversion for Personalized Text-to-Image Generation
by: Kim, Kunhee, et al.
Published: (2025)
by: Kim, Kunhee, et al.
Published: (2025)
SYNAuG: Exploiting Synthetic Data for Data Imbalance Problems
by: Ye-Bin, Moon, et al.
Published: (2023)
by: Ye-Bin, Moon, et al.
Published: (2023)
Boosting Weak Positives for Text Based Person Search
by: Modi, Akshay, et al.
Published: (2025)
by: Modi, Akshay, et al.
Published: (2025)
Steering Guidance for Personalized Text-to-Image Diffusion Models
by: Park, Sunghyun, et al.
Published: (2025)
by: Park, Sunghyun, et al.
Published: (2025)
Multi-level Cross-modal Alignment for Image Clustering
by: Qiu, Liping, et al.
Published: (2024)
by: Qiu, Liping, et al.
Published: (2024)
Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
Structured State-Space Regularization for Generation-Friendly Image Tokenization
by: Lee, Jinsung, et al.
Published: (2026)
by: Lee, Jinsung, et al.
Published: (2026)
Distilling Diffusion Models into Conditional GANs
by: Kang, Minguk, et al.
Published: (2024)
by: Kang, Minguk, et al.
Published: (2024)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
One-Cycle Structured Pruning via Stability-Driven Subnetwork Search
by: Ghimire, Deepak, et al.
Published: (2025)
by: Ghimire, Deepak, et al.
Published: (2025)
Multi-level and Multi-modal Action Anticipation
by: Kim, Seulgi, et al.
Published: (2025)
by: Kim, Seulgi, et al.
Published: (2025)
Retrieval-Augmented Score Distillation for Text-to-3D Generation
by: Seo, Junyoung, et al.
Published: (2024)
by: Seo, Junyoung, et al.
Published: (2024)
Identifiable Token Correspondence for World Models
by: Kim, Youngin, et al.
Published: (2026)
by: Kim, Youngin, et al.
Published: (2026)
RePL: Pseudo-label Refinement for Semi-supervised LiDAR Semantic Segmentation
by: Kwon, Donghyeon, et al.
Published: (2026)
by: Kwon, Donghyeon, et al.
Published: (2026)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
by: Wang, Hu, et al.
Published: (2023)
by: Wang, Hu, et al.
Published: (2023)
SCMM: Calibrating Cross-modal Representations for Text-Based Person Search
by: Liu, Jing, et al.
Published: (2023)
by: Liu, Jing, et al.
Published: (2023)
Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
by: Hong, Jinyung, et al.
Published: (2024)
by: Hong, Jinyung, et al.
Published: (2024)
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
by: Koishigarina, Darina, et al.
Published: (2025)
by: Koishigarina, Darina, et al.
Published: (2025)
Adversarial Robustification via Text-to-Image Diffusion Models
by: Choi, Daewon, et al.
Published: (2024)
by: Choi, Daewon, et al.
Published: (2024)
CorVS: Person Identification via Video Trajectory-Sensor Correspondence in a Real-World Warehouse
by: Kano, Kazuma, et al.
Published: (2025)
by: Kano, Kazuma, et al.
Published: (2025)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Generalizable Single-Source Cross-modality Medical Image Segmentation via Invariant Causal Mechanisms
by: Chen, Boqi, et al.
Published: (2024)
by: Chen, Boqi, et al.
Published: (2024)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion
by: Zhang, Yuguang, et al.
Published: (2025)
by: Zhang, Yuguang, et al.
Published: (2025)
Cross-Class Feature Augmentation for Class Incremental Learning
by: Kim, Taehoon, et al.
Published: (2023)
by: Kim, Taehoon, et al.
Published: (2023)
Similar Items
-
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
by: Park, Jicheol, et al.
Published: (2024) -
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025) -
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
by: Kim, Sungyeon, et al.
Published: (2024) -
Bootstrapping Top-down Information for Self-modulating Slot Attention
by: Kim, Dongwon, et al.
Published: (2024) -
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)