YotoR-You Only Transform One Representation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Villa, José Ignacio Díaz, Loncomilla, Patricio, Ruiz-del-Solar, Javier
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911894057517056
author Villa, José Ignacio Díaz
Loncomilla, Patricio
Ruiz-del-Solar, Javier
author_facet Villa, José Ignacio Díaz
Loncomilla, Patricio
Ruiz-del-Solar, Javier
contents This paper introduces YotoR (You Only Transform One Representation), a novel deep learning model for object detection that combines Swin Transformers and YoloR architectures. Transformers, a revolutionary technology in natural language processing, have also significantly impacted computer vision, offering the potential to enhance accuracy and computational efficiency. YotoR combines the robust Swin Transformer backbone with the YoloR neck and head. In our experiments, YotoR models TP5 and BP4 consistently outperform YoloR P6 and Swin Transformers in various evaluations, delivering improved object detection performance and faster inference speeds than Swin Transformer models. These results highlight the potential for further model combinations and improvements in real-time object detection with Transformers. The paper concludes by emphasizing the broader implications of YotoR, including its potential to enhance transformer-based models for image-related tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19629
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle YotoR-You Only Transform One Representation
Villa, José Ignacio Díaz
Loncomilla, Patricio
Ruiz-del-Solar, Javier
Computer Vision and Pattern Recognition
I.2.10
This paper introduces YotoR (You Only Transform One Representation), a novel deep learning model for object detection that combines Swin Transformers and YoloR architectures. Transformers, a revolutionary technology in natural language processing, have also significantly impacted computer vision, offering the potential to enhance accuracy and computational efficiency. YotoR combines the robust Swin Transformer backbone with the YoloR neck and head. In our experiments, YotoR models TP5 and BP4 consistently outperform YoloR P6 and Swin Transformers in various evaluations, delivering improved object detection performance and faster inference speeds than Swin Transformer models. These results highlight the potential for further model combinations and improvements in real-time object detection with Transformers. The paper concludes by emphasizing the broader implications of YotoR, including its potential to enhance transformer-based models for image-related tasks.
title YotoR-You Only Transform One Representation
topic Computer Vision and Pattern Recognition
I.2.10
url https://arxiv.org/abs/2405.19629