Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Lei, Li, Weihao, Dai, Pingyang, Chen, Jie, Cao, Liujuan, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification
by: Tan, Lei, et al.
Published: (2024)
by: Tan, Lei, et al.
Published: (2024)
Attention Disturbance and Dual-Path Constraint Network for Occluded Person Re-identification
by: Xia, Jiaer, et al.
Published: (2023)
by: Xia, Jiaer, et al.
Published: (2023)
Prompt Decoupling for Text-to-Image Person Re-identification
by: Li, Weihao, et al.
Published: (2024)
by: Li, Weihao, et al.
Published: (2024)
DPM++: Dynamic Masked Metric Learning for Occluded Person Re-identification
by: Tan, Lei, et al.
Published: (2026)
by: Tan, Lei, et al.
Published: (2026)
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
More Clear, More Flexible, More Precise: A Comprehensive Oriented Object Detection benchmark for UAV
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
Depth-Guided Semi-Supervised Instance Segmentation
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?
by: Jiang, Hongbo, et al.
Published: (2026)
by: Jiang, Hongbo, et al.
Published: (2026)
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
by: Huang, You, et al.
Published: (2025)
by: Huang, You, et al.
Published: (2025)
Unleashing MLLMs on the Edge: A Unified Framework for Cross-Modal ReID via Adaptive SVD Distillation
by: Jiang, Hongbo, et al.
Published: (2026)
by: Jiang, Hongbo, et al.
Published: (2026)
Purifying, Labeling, and Utilizing: A High-Quality Pipeline for Small Object Detection
by: Wang, Siwei, et al.
Published: (2025)
by: Wang, Siwei, et al.
Published: (2025)
Evolving, Not Training: Zero-Shot Reasoning Segmentation via Evolutionary Prompting
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
FocSAM: Delving Deeply into Focused Objects in Segmenting Anything
by: Huang, You, et al.
Published: (2024)
by: Huang, You, et al.
Published: (2024)
RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-identification
by: Tan, Lei, et al.
Published: (2024)
by: Tan, Lei, et al.
Published: (2024)
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval
by: Ma, Yiwei, et al.
Published: (2024)
by: Ma, Yiwei, et al.
Published: (2024)
Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text
by: Li, Xinyang, et al.
Published: (2024)
by: Li, Xinyang, et al.
Published: (2024)
Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
Dual3D: Efficient and Consistent Text-to-3D Generation with Dual-mode Multi-view Latent Diffusion
by: Li, Xinyang, et al.
Published: (2024)
by: Li, Xinyang, et al.
Published: (2024)
DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion
by: Qu, Yansong, et al.
Published: (2025)
by: Qu, Yansong, et al.
Published: (2025)
GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane
by: Qu, Yansong, et al.
Published: (2024)
by: Qu, Yansong, et al.
Published: (2024)
GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification
by: Li, Qiao, et al.
Published: (2025)
by: Li, Qiao, et al.
Published: (2025)
LAIP: Learning Local Alignment from Image-Phrase Modeling for Text-based Person Search
by: Wang, Haiguang, et al.
Published: (2024)
by: Wang, Haiguang, et al.
Published: (2024)
Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial Inspection
by: Lu, Yilin, et al.
Published: (2025)
by: Lu, Yilin, et al.
Published: (2025)
CutDiffusion: A Simple, Fast, Cheap, and Strong Diffusion Extrapolation Method
by: Lin, Mingbao, et al.
Published: (2024)
by: Lin, Mingbao, et al.
Published: (2024)
UniVST: A Unified Framework for Training-free Localized Video Style Transfer
by: Song, Quanjian, et al.
Published: (2024)
by: Song, Quanjian, et al.
Published: (2024)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
by: Li, Jiale, et al.
Published: (2025)
by: Li, Jiale, et al.
Published: (2025)
Drag Your Gaussian: Effective Drag-Based Editing with Score Distillation for 3D Gaussian Splatting
by: Qu, Yansong, et al.
Published: (2025)
by: Qu, Yansong, et al.
Published: (2025)
Semi-supervised Text-based Person Search
by: Gao, Daming, et al.
Published: (2024)
by: Gao, Daming, et al.
Published: (2024)
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
by: Ke, Shuyan, et al.
Published: (2026)
by: Ke, Shuyan, et al.
Published: (2026)
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
by: Lin, Weihuang, et al.
Published: (2025)
by: Lin, Weihuang, et al.
Published: (2025)
HRSAM: Efficient Interactive Segmentation in High-Resolution Images
by: Huang, You, et al.
Published: (2024)
by: Huang, You, et al.
Published: (2024)
HUWSOD: Holistic Self-training for Unified Weakly Supervised Object Detection
by: Cao, Liujuan, et al.
Published: (2024)
by: Cao, Liujuan, et al.
Published: (2024)
DiffusionFace: Towards a Comprehensive Dataset for Diffusion-Based Face Forgery Analysis
by: Chen, Zhongxi, et al.
Published: (2024)
by: Chen, Zhongxi, et al.
Published: (2024)
Breaking the Bias: Recalibrating the Attention of Industrial Anomaly Detection
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
Test-Time Computing for Referring Multimodal Large Language Models
by: Wu, Mingrui, et al.
Published: (2026)
by: Wu, Mingrui, et al.
Published: (2026)
Enhancing Visual Representation for Text-based Person Searching
by: Shen, Wei, et al.
Published: (2024)
by: Shen, Wei, et al.
Published: (2024)
Pseudo-Label Quality Decoupling and Correction for Semi-Supervised Instance Segmentation
by: Lin, Jianghang, et al.
Published: (2025)
by: Lin, Jianghang, et al.
Published: (2025)
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
by: Lin, Jianghang, et al.
Published: (2025)
by: Lin, Jianghang, et al.
Published: (2025)
Similar Items
-
PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification
by: Tan, Lei, et al.
Published: (2024) -
Attention Disturbance and Dual-Path Constraint Network for Occluded Person Re-identification
by: Xia, Jiaer, et al.
Published: (2023) -
Prompt Decoupling for Text-to-Image Person Re-identification
by: Li, Weihao, et al.
Published: (2024) -
DPM++: Dynamic Masked Metric Learning for Occluded Person Re-identification
by: Tan, Lei, et al.
Published: (2026) -
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
by: Sun, Zhen, et al.
Published: (2025)