PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912977179901952 |
|---|---|
| author | Kang, Zhengjian Zhuang, Jun Mo, Kangtong Chen, Qi Liu, Rui Zhang, Ye |
| author_facet | Kang, Zhengjian Zhuang, Jun Mo, Kangtong Chen, Qi Liu, Rui Zhang, Ye |
| contents | Detection Transformer (DETR) has redefined object detection by casting it as a set prediction task within an end-to-end framework. Despite its elegance, DETR and its variants still rely on fixed learnable queries and suffer from severe query utilization imbalance, which limits adaptability and leaves the model capacity underused. We propose PaQ-DETR (Pattern and Quality-Aware DETR), a unified framework that enhances both query adaptivity and supervision balance. It learns a compact set of shared latent patterns capturing global semantics and dynamically generates image-specific queries through content-conditioned weighting. In parallel, a quality-aware one-to-many assignment strategy adaptively selects positive samples based on localizatio-classification consistency, enriching supervision and promoting balanced query optimization. Experiments on COCO, CityScapes, and other benchmarks show consistent gains of 1.5%-4.2% mAP across DETR backbones, including ResNet and Swin-Transformer. Beyond accuracy improvement, our method provides interpretable insights into how dynamic patterns cluster semantically across object categories. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_06917 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection Kang, Zhengjian Zhuang, Jun Mo, Kangtong Chen, Qi Liu, Rui Zhang, Ye Computer Vision and Pattern Recognition Detection Transformer (DETR) has redefined object detection by casting it as a set prediction task within an end-to-end framework. Despite its elegance, DETR and its variants still rely on fixed learnable queries and suffer from severe query utilization imbalance, which limits adaptability and leaves the model capacity underused. We propose PaQ-DETR (Pattern and Quality-Aware DETR), a unified framework that enhances both query adaptivity and supervision balance. It learns a compact set of shared latent patterns capturing global semantics and dynamically generates image-specific queries through content-conditioned weighting. In parallel, a quality-aware one-to-many assignment strategy adaptively selects positive samples based on localizatio-classification consistency, enriching supervision and promoting balanced query optimization. Experiments on COCO, CityScapes, and other benchmarks show consistent gains of 1.5%-4.2% mAP across DETR backbones, including ResNet and Swin-Transformer. Beyond accuracy improvement, our method provides interpretable insights into how dynamic patterns cluster semantically across object categories. |
| title | PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.06917 |