Hyper-YOLO: When Visual Object Detection Meets Hypergraph Computation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Yifan, Huang, Jiangang, Du, Shaoyi, Ying, Shihui, Yong, Jun-Hai, Li, Yipeng, Ding, Guiguang, Ji, Rongrong, Gao, Yue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912074190290944
author Feng, Yifan
Huang, Jiangang
Du, Shaoyi
Ying, Shihui
Yong, Jun-Hai
Li, Yipeng
Ding, Guiguang
Ji, Rongrong
Gao, Yue
author_facet Feng, Yifan
Huang, Jiangang
Du, Shaoyi
Ying, Shihui
Yong, Jun-Hai
Li, Yipeng
Ding, Guiguang
Ji, Rongrong
Gao, Yue
contents We introduce Hyper-YOLO, a new object detection method that integrates hypergraph computations to capture the complex high-order correlations among visual features. Traditional YOLO models, while powerful, have limitations in their neck designs that restrict the integration of cross-level features and the exploitation of high-order feature interrelationships. To address these challenges, we propose the Hypergraph Computation Empowered Semantic Collecting and Scattering (HGC-SCS) framework, which transposes visual feature maps into a semantic space and constructs a hypergraph for high-order message propagation. This enables the model to acquire both semantic and structural information, advancing beyond conventional feature-focused learning. Hyper-YOLO incorporates the proposed Mixed Aggregation Network (MANet) in its backbone for enhanced feature extraction and introduces the Hypergraph-Based Cross-Level and Cross-Position Representation Network (HyperC2Net) in its neck. HyperC2Net operates across five scales and breaks free from traditional grid structures, allowing for sophisticated high-order interactions across levels and positions. This synergy of components positions Hyper-YOLO as a state-of-the-art architecture in various scale models, as evidenced by its superior performance on the COCO dataset. Specifically, Hyper-YOLO-N significantly outperforms the advanced YOLOv8-N and YOLOv9-T with 12\% $\text{AP}^{val}$ and 9\% $\text{AP}^{val}$ improvements. The source codes are at ttps://github.com/iMoonLab/Hyper-YOLO.
format Preprint
id arxiv_https___arxiv_org_abs_2408_04804
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hyper-YOLO: When Visual Object Detection Meets Hypergraph Computation
Feng, Yifan
Huang, Jiangang
Du, Shaoyi
Ying, Shihui
Yong, Jun-Hai
Li, Yipeng
Ding, Guiguang
Ji, Rongrong
Gao, Yue
Computer Vision and Pattern Recognition
We introduce Hyper-YOLO, a new object detection method that integrates hypergraph computations to capture the complex high-order correlations among visual features. Traditional YOLO models, while powerful, have limitations in their neck designs that restrict the integration of cross-level features and the exploitation of high-order feature interrelationships. To address these challenges, we propose the Hypergraph Computation Empowered Semantic Collecting and Scattering (HGC-SCS) framework, which transposes visual feature maps into a semantic space and constructs a hypergraph for high-order message propagation. This enables the model to acquire both semantic and structural information, advancing beyond conventional feature-focused learning. Hyper-YOLO incorporates the proposed Mixed Aggregation Network (MANet) in its backbone for enhanced feature extraction and introduces the Hypergraph-Based Cross-Level and Cross-Position Representation Network (HyperC2Net) in its neck. HyperC2Net operates across five scales and breaks free from traditional grid structures, allowing for sophisticated high-order interactions across levels and positions. This synergy of components positions Hyper-YOLO as a state-of-the-art architecture in various scale models, as evidenced by its superior performance on the COCO dataset. Specifically, Hyper-YOLO-N significantly outperforms the advanced YOLOv8-N and YOLOv9-T with 12\% $\text{AP}^{val}$ and 9\% $\text{AP}^{val}$ improvements. The source codes are at ttps://github.com/iMoonLab/Hyper-YOLO.
title Hyper-YOLO: When Visual Object Detection Meets Hypergraph Computation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.04804