OneNet: A Channel-Wise 1D Convolutional U-Net

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Byun, Sanghyun, Shah, Kayvan, Gang, Ayushi, Apton, Christopher, Song, Jacob, Chung, Woo Seong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913578664067072
author Byun, Sanghyun
Shah, Kayvan
Gang, Ayushi
Apton, Christopher
Song, Jacob
Chung, Woo Seong
author_facet Byun, Sanghyun
Shah, Kayvan
Gang, Ayushi
Apton, Christopher
Song, Jacob
Chung, Woo Seong
contents Many state-of-the-art computer vision architectures leverage U-Net for its adaptability and efficient feature extraction. However, the multi-resolution convolutional design often leads to significant computational demands, limiting deployment on edge devices. We present a streamlined alternative: a 1D convolutional encoder that retains accuracy while enhancing its suitability for edge applications. Our novel encoder architecture achieves semantic segmentation through channel-wise 1D convolutions combined with pixel-unshuffle operations. By incorporating PixelShuffle, known for improving accuracy in super-resolution tasks while reducing computational load, OneNet captures spatial relationships without requiring 2D convolutions, reducing parameters by up to 47%. Additionally, we explore a fully 1D encoder-decoder that achieves a 71% reduction in size, albeit with some accuracy loss. We benchmark our approach against U-Net variants across diverse mask-generation tasks, demonstrating that it preserves accuracy effectively. Although focused on image segmentation, this architecture is adaptable to other convolutional applications. Code for the project is available at https://github.com/shbyun080/OneNet .
format Preprint
id arxiv_https___arxiv_org_abs_2411_09838
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OneNet: A Channel-Wise 1D Convolutional U-Net
Byun, Sanghyun
Shah, Kayvan
Gang, Ayushi
Apton, Christopher
Song, Jacob
Chung, Woo Seong
Image and Video Processing
Computer Vision and Pattern Recognition
Many state-of-the-art computer vision architectures leverage U-Net for its adaptability and efficient feature extraction. However, the multi-resolution convolutional design often leads to significant computational demands, limiting deployment on edge devices. We present a streamlined alternative: a 1D convolutional encoder that retains accuracy while enhancing its suitability for edge applications. Our novel encoder architecture achieves semantic segmentation through channel-wise 1D convolutions combined with pixel-unshuffle operations. By incorporating PixelShuffle, known for improving accuracy in super-resolution tasks while reducing computational load, OneNet captures spatial relationships without requiring 2D convolutions, reducing parameters by up to 47%. Additionally, we explore a fully 1D encoder-decoder that achieves a 71% reduction in size, albeit with some accuracy loss. We benchmark our approach against U-Net variants across diverse mask-generation tasks, demonstrating that it preserves accuracy effectively. Although focused on image segmentation, this architecture is adaptable to other convolutional applications. Code for the project is available at https://github.com/shbyun080/OneNet .
title OneNet: A Channel-Wise 1D Convolutional U-Net
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.09838