Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Jingwei, Qiu, Zhixuan, Jin, Boyu, Siripong, Surasakdi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911141286903808
author Peng, Jingwei
Qiu, Zhixuan
Jin, Boyu
Siripong, Surasakdi
author_facet Peng, Jingwei
Qiu, Zhixuan
Jin, Boyu
Siripong, Surasakdi
contents Human action recognition often struggles with deep semantic understanding, complex contextual information, and fine-grained distinction, limitations that traditional methods frequently encounter when dealing with diverse video data. Inspired by the remarkable capabilities of large language models, this paper introduces LVLM-VAR, a novel framework that pioneers the application of pre-trained Vision-Language Large Models (LVLMs) to video action recognition, emphasizing enhanced accuracy and interpretability. Our method features a Video-to-Semantic-Tokens (VST) Module, which innovatively transforms raw video sequences into discrete, semantically and temporally consistent "semantic action tokens," effectively crafting an "action narrative" that is comprehensible to an LVLM. These tokens, combined with natural language instructions, are then processed by a LoRA-fine-tuned LVLM (e.g., LLaVA-13B) for robust action classification and semantic reasoning. LVLM-VAR not only achieves state-of-the-art or highly competitive performance on challenging benchmarks such as NTU RGB+D and NTU RGB+D 120, demonstrating significant improvements (e.g., 94.1% on NTU RGB+D X-Sub and 90.0% on NTU RGB+D 120 X-Set), but also substantially boosts model interpretability by generating natural language explanations for its predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05695
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
Peng, Jingwei
Qiu, Zhixuan
Jin, Boyu
Siripong, Surasakdi
Computer Vision and Pattern Recognition
Human action recognition often struggles with deep semantic understanding, complex contextual information, and fine-grained distinction, limitations that traditional methods frequently encounter when dealing with diverse video data. Inspired by the remarkable capabilities of large language models, this paper introduces LVLM-VAR, a novel framework that pioneers the application of pre-trained Vision-Language Large Models (LVLMs) to video action recognition, emphasizing enhanced accuracy and interpretability. Our method features a Video-to-Semantic-Tokens (VST) Module, which innovatively transforms raw video sequences into discrete, semantically and temporally consistent "semantic action tokens," effectively crafting an "action narrative" that is comprehensible to an LVLM. These tokens, combined with natural language instructions, are then processed by a LoRA-fine-tuned LVLM (e.g., LLaVA-13B) for robust action classification and semantic reasoning. LVLM-VAR not only achieves state-of-the-art or highly competitive performance on challenging benchmarks such as NTU RGB+D and NTU RGB+D 120, demonstrating significant improvements (e.g., 94.1% on NTU RGB+D X-Sub and 90.0% on NTU RGB+D 120 X-Set), but also substantially boosts model interpretability by generating natural language explanations for its predictions.
title Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.05695