Heatmap Pooling Network for Action Recognition from RGB Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Mengyuan, Liu, Jinfu, Jiang, Yongkang, He, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918229984673792
author Liu, Mengyuan
Liu, Jinfu
Jiang, Yongkang
He, Bin
author_facet Liu, Mengyuan
Liu, Jinfu
Jiang, Yongkang
He, Bin
contents Human action recognition (HAR) in videos has garnered widespread attention due to the rich information in RGB videos. Nevertheless, existing methods for extracting deep features from RGB videos face challenges such as information redundancy, susceptibility to noise and high storage costs. To address these issues and fully harness the useful information in videos, we propose a novel heatmap pooling network (HP-Net) for action recognition from videos, which extracts information-rich, robust and concise pooled features of the human body in videos through a feedback pooling module. The extracted pooled features demonstrate obvious performance advantages over the previously obtained pose data and heatmap features from videos. In addition, we design a spatial-motion co-learning module and a text refinement modulation module to integrate the extracted pooled features with other multimodal data, enabling more robust action recognition. Extensive experiments on several benchmarks namely NTU RGB+D 60, NTU RGB+D 120, Toyota-Smarthome and UAV-Human consistently verify the effectiveness of our HP-Net, which outperforms the existing human action recognition methods. Our code is publicly available at: https://github.com/liujf69/HPNet-Action.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03837
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Heatmap Pooling Network for Action Recognition from RGB Videos
Liu, Mengyuan
Liu, Jinfu
Jiang, Yongkang
He, Bin
Computer Vision and Pattern Recognition
Human action recognition (HAR) in videos has garnered widespread attention due to the rich information in RGB videos. Nevertheless, existing methods for extracting deep features from RGB videos face challenges such as information redundancy, susceptibility to noise and high storage costs. To address these issues and fully harness the useful information in videos, we propose a novel heatmap pooling network (HP-Net) for action recognition from videos, which extracts information-rich, robust and concise pooled features of the human body in videos through a feedback pooling module. The extracted pooled features demonstrate obvious performance advantages over the previously obtained pose data and heatmap features from videos. In addition, we design a spatial-motion co-learning module and a text refinement modulation module to integrate the extracted pooled features with other multimodal data, enabling more robust action recognition. Extensive experiments on several benchmarks namely NTU RGB+D 60, NTU RGB+D 120, Toyota-Smarthome and UAV-Human consistently verify the effectiveness of our HP-Net, which outperforms the existing human action recognition methods. Our code is publicly available at: https://github.com/liujf69/HPNet-Action.
title Heatmap Pooling Network for Action Recognition from RGB Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.03837