FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Toupas, Petros, Bouganis, Christos-Savvas, Tzovaras, Dimitrios
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917603291693056
author Toupas, Petros
Bouganis, Christos-Savvas
Tzovaras, Dimitrios
author_facet Toupas, Petros
Bouganis, Christos-Savvas
Tzovaras, Dimitrios
contents 3D Convolutional Neural Networks are gaining increasing attention from researchers and practitioners and have found applications in many domains, such as surveillance systems, autonomous vehicles, human monitoring systems, and video retrieval. However, their widespread adoption is hindered by their high computational and memory requirements, especially when resource-constrained systems are targeted. This paper addresses the problem of mapping X3D, a state-of-the-art model in Human Action Recognition that achieves accuracy of 95.5\% in the UCF101 benchmark, onto any FPGA device. The proposed toolflow generates an optimised stream-based hardware system, taking into account the available resources and off-chip memory characteristics of the FPGA device. The generated designs push further the current performance-accuracy pareto front, and enable for the first time the targeting of such complex model architectures for the Human Action Recognition task.
format Preprint
id arxiv_https___arxiv_org_abs_2305_18479
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
Toupas, Petros
Bouganis, Christos-Savvas
Tzovaras, Dimitrios
Computer Vision and Pattern Recognition
Artificial Intelligence
Hardware Architecture
Machine Learning
3D Convolutional Neural Networks are gaining increasing attention from researchers and practitioners and have found applications in many domains, such as surveillance systems, autonomous vehicles, human monitoring systems, and video retrieval. However, their widespread adoption is hindered by their high computational and memory requirements, especially when resource-constrained systems are targeted. This paper addresses the problem of mapping X3D, a state-of-the-art model in Human Action Recognition that achieves accuracy of 95.5\% in the UCF101 benchmark, onto any FPGA device. The proposed toolflow generates an optimised stream-based hardware system, taking into account the available resources and off-chip memory characteristics of the FPGA device. The generated designs push further the current performance-accuracy pareto front, and enable for the first time the targeting of such complex model architectures for the Human Action Recognition task.
title FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2305.18479