CLEAR: Cross-Transformers with Pre-trained Language Model is All you need for Person Attribute Recognition and Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Bui, Doanh C., Le, Thinh V., Ngo, Ba Hung, Choi, Tae Jong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiGDA: Hierarchical Graph of Nodes to Learn Local-to-Global Topology for Semi-Supervised Domain Adaptation
by: Ngo, Ba Hung, et al.
Published: (2024)
by: Ngo, Ba Hung, et al.
Published: (2024)
Welcome New Doctor: Continual Learning with Expert Consultation and Autoregressive Inference for Whole Slide Image Analysis
by: Bui, Doanh Cao, et al.
Published: (2025)
by: Bui, Doanh Cao, et al.
Published: (2025)
MECFormer: Multi-task Whole Slide Image Classification with Expert Consultation Network
by: Bui, Doanh C., et al.
Published: (2024)
by: Bui, Doanh C., et al.
Published: (2024)
QuIIL at T3 challenge: Towards Automation in Life-Saving Intervention Procedures from First-Person View
by: Vuong, Trinh T. L., et al.
Published: (2024)
by: Vuong, Trinh T. L., et al.
Published: (2024)
MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
by: Bui, Doanh C., et al.
Published: (2025)
by: Bui, Doanh C., et al.
Published: (2025)
FALFormer: Feature-aware Landmarks self-attention for Whole-slide Image Classification
by: Bui, Doanh C., et al.
Published: (2024)
by: Bui, Doanh C., et al.
Published: (2024)
Learning CNN on ViT: A Hybrid Model to Explicitly Class-specific Boundaries for Domain Adaptation
by: Ngo, Ba Hung, et al.
Published: (2024)
by: Ngo, Ba Hung, et al.
Published: (2024)
CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
Open-Attribute Recognition for Person Retrieval: Finding People Through Distinctive and Novel Attributes
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
OCR is All you need: Importing Multi-Modality into Image-based Defect Detection System
by: Hsu, Chih-Chung, et al.
Published: (2024)
by: Hsu, Chih-Chung, et al.
Published: (2024)
Are CLIP features all you need for Universal Synthetic Image Origin Attribution?
by: Cioni, Dario, et al.
Published: (2024)
by: Cioni, Dario, et al.
Published: (2024)
BTS-rPPG: Orthogonal Butterfly Temporal Shifting for Remote Photoplethysmography
by: Nguyen, Ba-Thinh, et al.
Published: (2026)
by: Nguyen, Ba-Thinh, et al.
Published: (2026)
HOT: Harmonic-Constrained Optimal Transport for Remote Photoplethysmography Domain Adaptation
by: Nguyen, Ba-Thinh, et al.
Published: (2026)
by: Nguyen, Ba-Thinh, et al.
Published: (2026)
Label-Efficient Cross-Modality Generalization for Liver Segmentation in Multi-Phase MRI
by: Bui-Tran, Quang-Khai, et al.
Published: (2025)
by: Bui-Tran, Quang-Khai, et al.
Published: (2025)
Cross-video Identity Correlating for Person Re-identification Pre-training
by: Zuo, Jialong, et al.
Published: (2024)
by: Zuo, Jialong, et al.
Published: (2024)
Not All Starting Points Are Equal: Pre-trained Priors and Their Outsized Impact on Person Identification
by: Metz, Thomas M., et al.
Published: (2025)
by: Metz, Thomas M., et al.
Published: (2025)
PLIP: Language-Image Pre-training for Person Representation Learning
by: Zuo, Jialong, et al.
Published: (2023)
by: Zuo, Jialong, et al.
Published: (2023)
Reperio-rPPG: Relational Temporal Graph Neural Networks for Periodicity Learning in Remote Physiological Measurement
by: Nguyen, Ba-Thinh, et al.
Published: (2025)
by: Nguyen, Ba-Thinh, et al.
Published: (2025)
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
by: Huang, Jiale, et al.
Published: (2024)
by: Huang, Jiale, et al.
Published: (2024)
Logos as a Well-Tempered Pre-train for Sign Language Recognition
by: Ovodov, Ilya, et al.
Published: (2025)
by: Ovodov, Ilya, et al.
Published: (2025)
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
by: Zang, Yuan, et al.
Published: (2024)
by: Zang, Yuan, et al.
Published: (2024)
Long-Tailed Recognition on Binary Networks by Calibrating A Pre-trained Model
by: Kim, Jihun, et al.
Published: (2024)
by: Kim, Jihun, et al.
Published: (2024)
Spatio-Temporal Side Tuning Pre-trained Foundation Models for Video-based Pedestrian Attribute Recognition
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
ZeroSlide: Is Zero-Shot Classification Adequate for Lifelong Learning in Whole-Slide Image Analysis in the Era of Pathology Vision-Language Foundation Models?
by: Bui, Doanh C., et al.
Published: (2025)
by: Bui, Doanh C., et al.
Published: (2025)
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
by: Le, Kha Nhat, et al.
Published: (2024)
by: Le, Kha Nhat, et al.
Published: (2024)
CATFace: Cross-Attribute-Guided Transformer with Self-Attention Distillation for Low-Quality Face Recognition
by: Talemi, Niloufar Alipour, et al.
Published: (2024)
by: Talemi, Niloufar Alipour, et al.
Published: (2024)
Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners
by: Park, Keon-Hee, et al.
Published: (2024)
by: Park, Keon-Hee, et al.
Published: (2024)
Lifelong Whole Slide Image Analysis: Online Vision-Language Adaptation and Past-to-Present Gradient Distillation
by: Bui, Doanh C., et al.
Published: (2025)
by: Bui, Doanh C., et al.
Published: (2025)
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
by: Mahanta, Cristina, et al.
Published: (2025)
by: Mahanta, Cristina, et al.
Published: (2025)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
POA: Pre-training Once for Models of All Sizes
by: Zhang, Yingying, et al.
Published: (2024)
by: Zhang, Yingying, et al.
Published: (2024)
Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging
by: Han, Zihao, et al.
Published: (2025)
by: Han, Zihao, et al.
Published: (2025)
Swap Path Network for Robust Person Search Pre-training
by: Jaffe, Lucas, et al.
Published: (2024)
by: Jaffe, Lucas, et al.
Published: (2024)
Pre-training Enables Extraordinary All-optical Image Denoising
by: Lv, Xudong, et al.
Published: (2026)
by: Lv, Xudong, et al.
Published: (2026)
Pre-training for Action Recognition with Automatically Generated Fractal Datasets
by: Svyezhentsev, Davyd, et al.
Published: (2024)
by: Svyezhentsev, Davyd, et al.
Published: (2024)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
All you need is Trotter
by: Rendon, Gumaro
Published: (2023)
by: Rendon, Gumaro
Published: (2023)
Entrepreneurial orientation, entrepreneurial resources, and entrepreneurial success: The need for further exploration
by: Cong Doanh Duong
Published: (2022)
by: Cong Doanh Duong
Published: (2022)
CLEAR: Causal Learning Framework For Robust Histopathology Tumor Detection Under Out-Of-Distribution Shifts
by: Thi, Kieu-Anh Truong, et al.
Published: (2025)
by: Thi, Kieu-Anh Truong, et al.
Published: (2025)
Similar Items
-
HiGDA: Hierarchical Graph of Nodes to Learn Local-to-Global Topology for Semi-Supervised Domain Adaptation
by: Ngo, Ba Hung, et al.
Published: (2024) -
Welcome New Doctor: Continual Learning with Expert Consultation and Autoregressive Inference for Whole Slide Image Analysis
by: Bui, Doanh Cao, et al.
Published: (2025) -
MECFormer: Multi-task Whole Slide Image Classification with Expert Consultation Network
by: Bui, Doanh C., et al.
Published: (2024) -
QuIIL at T3 challenge: Towards Automation in Life-Saving Intervention Procedures from First-Person View
by: Vuong, Trinh T. L., et al.
Published: (2024) -
MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
by: Bui, Doanh C., et al.
Published: (2025)