YOLOv8 to YOLO11: A Comprehensive Architecture In-depth Comparative Review

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hidayatullah, Priyanto, Syakrani, Nurjannah, Sholahuddin, Muhammad Rizqi, Gelar, Trisna, Tubagus, Refdinal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914527859179520
author Hidayatullah, Priyanto
Syakrani, Nurjannah
Sholahuddin, Muhammad Rizqi
Gelar, Trisna
Tubagus, Refdinal
author_facet Hidayatullah, Priyanto
Syakrani, Nurjannah
Sholahuddin, Muhammad Rizqi
Gelar, Trisna
Tubagus, Refdinal
contents Note: This is a preliminary version of the manuscript. The final, peer-reviewed, and substantially revised version has been published in Jurnal RESTI. Readers are encouraged to access and cite the published version: DOI: https://doi.org/10.29207/resti.v10i2.6598 In the field of deep learning-based computer vision, YOLO is revolutionary. With respect to deep learning models, YOLO is also the one that is evolving the most rapidly. Unfortunately, not every YOLO model possesses scholarly publications. Moreover, there exists a YOLO model that lacks a publicly accessible official architectural diagram. Naturally, this engenders challenges, such as complicating the understanding of how the model operates in practice. Furthermore, the review articles that are presently available do not investigate the specifics of each model. The objective of this study is to present a comprehensive and in-depth architecture comparison of the four most recent YOLO models, specifically YOLOv8 through YOLO11, thereby enabling readers to quickly grasp not only how each model functions, but also the distinctions between them. To analyze each YOLO version's architecture, we meticulously examined the relevant academic papers, documentation, and scrutinized the source code. The analysis reveals that while each version of YOLO has improvements in architecture and feature extraction, certain blocks remain unchanged. The lack of scholarly publications and official diagrams presents challenges for understanding the model's functionality and future enhancement. Future developers are encouraged to provide these resources.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle YOLOv8 to YOLO11: A Comprehensive Architecture In-depth Comparative Review
Hidayatullah, Priyanto
Syakrani, Nurjannah
Sholahuddin, Muhammad Rizqi
Gelar, Trisna
Tubagus, Refdinal
Computer Vision and Pattern Recognition
Artificial Intelligence
Note: This is a preliminary version of the manuscript. The final, peer-reviewed, and substantially revised version has been published in Jurnal RESTI. Readers are encouraged to access and cite the published version: DOI: https://doi.org/10.29207/resti.v10i2.6598 In the field of deep learning-based computer vision, YOLO is revolutionary. With respect to deep learning models, YOLO is also the one that is evolving the most rapidly. Unfortunately, not every YOLO model possesses scholarly publications. Moreover, there exists a YOLO model that lacks a publicly accessible official architectural diagram. Naturally, this engenders challenges, such as complicating the understanding of how the model operates in practice. Furthermore, the review articles that are presently available do not investigate the specifics of each model. The objective of this study is to present a comprehensive and in-depth architecture comparison of the four most recent YOLO models, specifically YOLOv8 through YOLO11, thereby enabling readers to quickly grasp not only how each model functions, but also the distinctions between them. To analyze each YOLO version's architecture, we meticulously examined the relevant academic papers, documentation, and scrutinized the source code. The analysis reveals that while each version of YOLO has improvements in architecture and feature extraction, certain blocks remain unchanged. The lack of scholarly publications and official diagrams presents challenges for understanding the model's functionality and future enhancement. Future developers are encouraged to provide these resources.
title YOLOv8 to YOLO11: A Comprehensive Architecture In-depth Comparative Review
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2501.13400