VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Sellam, Abdellah Zakaria, Bekhouche, Salah Eddine, Dornaika, Fadi, Distante, Cosimo, Hadid, Abdenour |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
by: Sellam, Abdellah Zakaria, et al.
Published: (2025)
by: Sellam, Abdellah Zakaria, et al.
Published: (2025)
VP-Hype: A Hybrid Mamba-Transformer Framework with Visual-Textual Prompting for Hyperspectral Image Classification
by: Sellam, Abdellah Zakaria, et al.
Published: (2026)
by: Sellam, Abdellah Zakaria, et al.
Published: (2026)
RF-HiT: Rectified Flow Hierarchical Transformer for General Medical Image Segmentation
by: Djouama, Ahmed Marouane, et al.
Published: (2026)
by: Djouama, Ahmed Marouane, et al.
Published: (2026)
SegDT: A Diffusion Transformer-Based Segmentation Model for Medical Imaging
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
by: Sellam, Abdellah Zakaria, et al.
Published: (2025)
by: Sellam, Abdellah Zakaria, et al.
Published: (2025)
LoLA-SpecViT: Local Attention SwiGLU Vision Transformer with LoRA for Hyperspectral Imaging
by: Zidi, Fadi Abdeladhim, et al.
Published: (2025)
by: Zidi, Fadi Abdeladhim, et al.
Published: (2025)
CVPD at QIAS 2025 Shared Task: An Efficient Encoder-Based Approach for Islamic Inheritance Reasoning
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
Integrating ConvNeXt and Vision Transformers for Enhancing Facial Age Estimation
by: Maroun, Gaby, et al.
Published: (2025)
by: Maroun, Gaby, et al.
Published: (2025)
Decoding Matters: Efficient Mamba-Based Decoder with Distribution-Aware Deep Supervision for Medical Image Segmentation
by: Bougourzi, Fares, et al.
Published: (2026)
by: Bougourzi, Fares, et al.
Published: (2026)
Conflict-Aware Multimodal Fusion for Ambivalence and Hesitancy Recognition
by: Bekhouche, Salah Eddine, et al.
Published: (2026)
by: Bekhouche, Salah Eddine, et al.
Published: (2026)
SPARK-IL: Spectral Retrieval-Augmented RAG for Knowledge-driven Deepfake Detection via Incremental Learning
by: Eutamene, Hessen Bougueffa, et al.
Published: (2026)
by: Eutamene, Hessen Bougueffa, et al.
Published: (2026)
UniPAR: A Unified Framework for Pedestrian Attribute Recognition
by: Xu, Minghe, et al.
Published: (2026)
by: Xu, Minghe, et al.
Published: (2026)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling
by: Zhu, William Yicheng, et al.
Published: (2024)
by: Zhu, William Yicheng, et al.
Published: (2024)
Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
by: Bekhouche, Salah Eddine, et al.
Published: (2025)
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
D-TrAttUnet: Toward Hybrid CNN-Transformer Architecture for Generic and Subtle Segmentation in Medical Images
by: Bougourzi, Fares, et al.
Published: (2024)
by: Bougourzi, Fares, et al.
Published: (2024)
DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
by: Sufian, Abu, et al.
Published: (2025)
by: Sufian, Abu, et al.
Published: (2025)
SNN-PAR: Energy Efficient Pedestrian Attribute Recognition via Spiking Neural Networks
by: Wang, Haiyang, et al.
Published: (2024)
by: Wang, Haiyang, et al.
Published: (2024)
An Empirical Study of Mamba-based Pedestrian Attribute Recognition
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
FIDAVL: Fake Image Detection and Attribution using Vision-Language Model
by: Keita, Mamadou, et al.
Published: (2024)
by: Keita, Mamadou, et al.
Published: (2024)
Pedestrian Attribute Recognition via Hierarchical Cross-Modality HyperGraph Learning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Cross-Modal Mapping and Dual-Branch Reconstruction for 2D-3D Multimodal Industrial Anomaly Detection
by: Daci, Radia, et al.
Published: (2026)
by: Daci, Radia, et al.
Published: (2026)
Pedestrian Attribute Recognition: A New Benchmark Dataset and A Large Language Model Augmented Framework
by: Jin, Jiandong, et al.
Published: (2024)
by: Jin, Jiandong, et al.
Published: (2024)
Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute Recognition
by: Kong, Weizhe, et al.
Published: (2025)
by: Kong, Weizhe, et al.
Published: (2025)
Local and Global Context-and-Object-part-Aware Superpixel-based Data Augmentation for Deep Visual Recognition
by: Dornaika, Fadi, et al.
Published: (2025)
by: Dornaika, Fadi, et al.
Published: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
by: Park, Junghyun, et al.
Published: (2025)
by: Park, Junghyun, et al.
Published: (2025)
Recent Advances in Medical Imaging Segmentation: A Survey
by: Bougourzi, Fares, et al.
Published: (2025)
by: Bougourzi, Fares, et al.
Published: (2025)
RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Real-Time Indoor Object Detection based on hybrid CNN-Transformer Approach
by: Laidoudi, Salah Eddine, et al.
Published: (2024)
by: Laidoudi, Salah Eddine, et al.
Published: (2024)
PFM-VEPAR: Prompting Foundation Models for RGB-Event Camera based Pedestrian Attribute Recognition
by: Xu, Minghe, et al.
Published: (2026)
by: Xu, Minghe, et al.
Published: (2026)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026)
by: Mishra, Naman, et al.
Published: (2026)
PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition
by: Saadi, Ibtissam, et al.
Published: (2025)
by: Saadi, Ibtissam, et al.
Published: (2025)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
by: He, Chiyuan, et al.
Published: (2025)
by: He, Chiyuan, et al.
Published: (2025)
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
by: Liu, Hanqing, et al.
Published: (2026)
by: Liu, Hanqing, et al.
Published: (2026)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
by: Xu, Wanting, et al.
Published: (2024)
by: Xu, Wanting, et al.
Published: (2024)
Similar Items
-
C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
by: Sellam, Abdellah Zakaria, et al.
Published: (2025) -
VP-Hype: A Hybrid Mamba-Transformer Framework with Visual-Textual Prompting for Hyperspectral Image Classification
by: Sellam, Abdellah Zakaria, et al.
Published: (2026) -
RF-HiT: Rectified Flow Hierarchical Transformer for General Medical Image Segmentation
by: Djouama, Ahmed Marouane, et al.
Published: (2026) -
SegDT: A Diffusion Transformer-Based Segmentation Model for Medical Imaging
by: Bekhouche, Salah Eddine, et al.
Published: (2025) -
Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
by: Sellam, Abdellah Zakaria, et al.
Published: (2025)