RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916335033778176 |
|---|---|
| author | Lv, Wenyu Zhao, Yian Chang, Qinyao Huang, Kui Wang, Guanzhong Liu, Yi |
| author_facet | Lv, Wenyu Zhao, Yian Chang, Qinyao Huang, Kui Wang, Guanzhong Liu, Yi |
| contents | In this report, we present RT-DETRv2, an improved Real-Time DEtection TRansformer (RT-DETR). RT-DETRv2 builds upon the previous state-of-the-art real-time detector, RT-DETR, and opens up a set of bag-of-freebies for flexibility and practicality, as well as optimizing the training strategy to achieve enhanced performance. To improve the flexibility, we suggest setting a distinct number of sampling points for features at different scales in the deformable attention to achieve selective multi-scale feature extraction by the decoder. To enhance practicality, we propose an optional discrete sampling operator to replace the grid_sample operator that is specific to RT-DETR compared to YOLOs. This removes the deployment constraints typically associated with DETRs. For the training strategy, we propose dynamic data augmentation and scale-adaptive hyperparameters customization to improve performance without loss of speed. Source code and pre-trained models will be available at https://github.com/lyuwenyu/RT-DETR. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_17140 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer Lv, Wenyu Zhao, Yian Chang, Qinyao Huang, Kui Wang, Guanzhong Liu, Yi Computer Vision and Pattern Recognition In this report, we present RT-DETRv2, an improved Real-Time DEtection TRansformer (RT-DETR). RT-DETRv2 builds upon the previous state-of-the-art real-time detector, RT-DETR, and opens up a set of bag-of-freebies for flexibility and practicality, as well as optimizing the training strategy to achieve enhanced performance. To improve the flexibility, we suggest setting a distinct number of sampling points for features at different scales in the deformable attention to achieve selective multi-scale feature extraction by the decoder. To enhance practicality, we propose an optional discrete sampling operator to replace the grid_sample operator that is specific to RT-DETR compared to YOLOs. This removes the deployment constraints typically associated with DETRs. For the training strategy, we propose dynamic data augmentation and scale-adaptive hyperparameters customization to improve performance without loss of speed. Source code and pre-trained models will be available at https://github.com/lyuwenyu/RT-DETR. |
| title | RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2407.17140 |