VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915696339845120 |
|---|---|
| author | Sellam, Abdellah Zakaria Bekhouche, Salah Eddine Dornaika, Fadi Distante, Cosimo Hadid, Abdenour |
| author_facet | Sellam, Abdellah Zakaria Bekhouche, Salah Eddine Dornaika, Fadi Distante, Cosimo Hadid, Abdenour |
| contents | Pedestrian Attribute Recognition (PAR) involves predicting fine-grained attributes such as clothing color, gender, and accessories from pedestrian imagery, yet is hindered by severe class imbalance, intricate attribute co-dependencies, and domain shifts. We introduce VLM-PAR, a modular vision-language framework built on frozen SigLIP 2 multilingual encoders. By first aligning image and prompt embeddings via refining visual features through a compact cross-attention fusion, VLM-PAR achieves significant accuracy improvement on the highly imbalanced PA100K benchmark, setting a new state-of-the-art performance, while also delivering significant gains in mean accuracy across PETA and Market-1501 benchmarks. These results underscore the efficacy of integrating large-scale vision-language pretraining with targeted cross-modal refinement to overcome imbalance and generalization challenges in PAR. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_22217 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition Sellam, Abdellah Zakaria Bekhouche, Salah Eddine Dornaika, Fadi Distante, Cosimo Hadid, Abdenour Computer Vision and Pattern Recognition Artificial Intelligence Pedestrian Attribute Recognition (PAR) involves predicting fine-grained attributes such as clothing color, gender, and accessories from pedestrian imagery, yet is hindered by severe class imbalance, intricate attribute co-dependencies, and domain shifts. We introduce VLM-PAR, a modular vision-language framework built on frozen SigLIP 2 multilingual encoders. By first aligning image and prompt embeddings via refining visual features through a compact cross-attention fusion, VLM-PAR achieves significant accuracy improvement on the highly imbalanced PA100K benchmark, setting a new state-of-the-art performance, while also delivering significant gains in mean accuracy across PETA and Market-1501 benchmarks. These results underscore the efficacy of integrating large-scale vision-language pretraining with targeted cross-modal refinement to overcome imbalance and generalization challenges in PAR. |
| title | VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2512.22217 |