Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qammaz, Ammar, Vasilikopoulos, Nikolaos, Oikonomidis, Iason, Argyros, Antonis A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910699868913664
author Qammaz, Ammar
Vasilikopoulos, Nikolaos
Oikonomidis, Iason
Argyros, Antonis A.
author_facet Qammaz, Ammar
Vasilikopoulos, Nikolaos
Oikonomidis, Iason
Argyros, Antonis A.
contents We present Y-MAP-Net, a Y-shaped neural network architecture designed for real-time multi-task learning on RGB images. Y-MAP-Net, simultaneously predicts depth, surface normals, human pose, semantic segmentation and generates multi-label captions, all from a single network evaluation. To achieve this, we adopt a multi-teacher, single-student training paradigm, where task-specific foundation models supervise the network's learning, enabling it to distill their capabilities into a lightweight architecture suitable for real-time applications. Y-MAP-Net, exhibits strong generalization, simplicity and computational efficiency, making it ideal for robotics and other practical scenarios. To support future research, we will release our code publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10334
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images
Qammaz, Ammar
Vasilikopoulos, Nikolaos
Oikonomidis, Iason
Argyros, Antonis A.
Computer Vision and Pattern Recognition
We present Y-MAP-Net, a Y-shaped neural network architecture designed for real-time multi-task learning on RGB images. Y-MAP-Net, simultaneously predicts depth, surface normals, human pose, semantic segmentation and generates multi-label captions, all from a single network evaluation. To achieve this, we adopt a multi-teacher, single-student training paradigm, where task-specific foundation models supervise the network's learning, enabling it to distill their capabilities into a lightweight architecture suitable for real-time applications. Y-MAP-Net, exhibits strong generalization, simplicity and computational efficiency, making it ideal for robotics and other practical scenarios. To support future research, we will release our code publicly.
title Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.10334