VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sellam, Abdellah Zakaria, Bekhouche, Salah Eddine, Dornaika, Fadi, Distante, Cosimo, Hadid, Abdenour
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915696339845120
author Sellam, Abdellah Zakaria
Bekhouche, Salah Eddine
Dornaika, Fadi
Distante, Cosimo
Hadid, Abdenour
author_facet Sellam, Abdellah Zakaria
Bekhouche, Salah Eddine
Dornaika, Fadi
Distante, Cosimo
Hadid, Abdenour
contents Pedestrian Attribute Recognition (PAR) involves predicting fine-grained attributes such as clothing color, gender, and accessories from pedestrian imagery, yet is hindered by severe class imbalance, intricate attribute co-dependencies, and domain shifts. We introduce VLM-PAR, a modular vision-language framework built on frozen SigLIP 2 multilingual encoders. By first aligning image and prompt embeddings via refining visual features through a compact cross-attention fusion, VLM-PAR achieves significant accuracy improvement on the highly imbalanced PA100K benchmark, setting a new state-of-the-art performance, while also delivering significant gains in mean accuracy across PETA and Market-1501 benchmarks. These results underscore the efficacy of integrating large-scale vision-language pretraining with targeted cross-modal refinement to overcome imbalance and generalization challenges in PAR.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
Sellam, Abdellah Zakaria
Bekhouche, Salah Eddine
Dornaika, Fadi
Distante, Cosimo
Hadid, Abdenour
Computer Vision and Pattern Recognition
Artificial Intelligence
Pedestrian Attribute Recognition (PAR) involves predicting fine-grained attributes such as clothing color, gender, and accessories from pedestrian imagery, yet is hindered by severe class imbalance, intricate attribute co-dependencies, and domain shifts. We introduce VLM-PAR, a modular vision-language framework built on frozen SigLIP 2 multilingual encoders. By first aligning image and prompt embeddings via refining visual features through a compact cross-attention fusion, VLM-PAR achieves significant accuracy improvement on the highly imbalanced PA100K benchmark, setting a new state-of-the-art performance, while also delivering significant gains in mean accuracy across PETA and Market-1501 benchmarks. These results underscore the efficacy of integrating large-scale vision-language pretraining with targeted cross-modal refinement to overcome imbalance and generalization challenges in PAR.
title VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.22217