Research on a hybrid LSTM-CNN-Attention model for text-based web content classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuz, Mykola, Lazarovych, Ihor, Kozlenko, Mykola, Pikuliak, Mykola, Kvasniuk, Andrii
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918263996284928
author Kuz, Mykola
Lazarovych, Ihor
Kozlenko, Mykola
Pikuliak, Mykola
Kvasniuk, Andrii
author_facet Kuz, Mykola
Lazarovych, Ihor
Kozlenko, Mykola
Pikuliak, Mykola
Kvasniuk, Andrii
contents This study presents a hybrid deep learning architecture that integrates LSTM, CNN, and an Attention mechanism to enhance the classification of web content based on text. Pretrained GloVe embeddings are used to represent words as dense vectors that preserve semantic similarity. The CNN layer extracts local n-gram patterns and lexical features, while the LSTM layer models long-range dependencies and sequential structure. The integrated Attention mechanism enables the model to focus selectively on the most informative parts of the input sequence. A 5-fold cross-validation setup was used to assess the robustness and generalizability of the proposed solution. Experimental results show that the hybrid LSTM-CNN-Attention model achieved outstanding performance, with an accuracy of 0.98, precision of 0.94, recall of 0.92, and F1-score of 0.93. These results surpass the performance of baseline models based solely on CNNs, LSTMs, or transformer-based classifiers such as BERT. The combination of neural network components enabled the model to effectively capture both fine-grained text structures and broader semantic context. Furthermore, the use of GloVe embeddings provided an efficient and effective representation of textual data, making the model suitable for integration into systems with real-time or near-real-time requirements. The proposed hybrid architecture demonstrates high effectiveness in text-based web content classification, particularly in tasks requiring both syntactic feature extraction and semantic interpretation. By combining presented mechanisms, the model addresses the limitations of individual architectures and achieves improved generalization. These findings support the broader use of hybrid deep learning approaches in NLP applications, especially where complex, unstructured textual data must be processed and classified with high reliability.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
Kuz, Mykola
Lazarovych, Ihor
Kozlenko, Mykola
Pikuliak, Mykola
Kvasniuk, Andrii
Computation and Language
Machine Learning
68T50 (Primary) 68T07 (Secondary)
I.2.7; I.2.6
This study presents a hybrid deep learning architecture that integrates LSTM, CNN, and an Attention mechanism to enhance the classification of web content based on text. Pretrained GloVe embeddings are used to represent words as dense vectors that preserve semantic similarity. The CNN layer extracts local n-gram patterns and lexical features, while the LSTM layer models long-range dependencies and sequential structure. The integrated Attention mechanism enables the model to focus selectively on the most informative parts of the input sequence. A 5-fold cross-validation setup was used to assess the robustness and generalizability of the proposed solution. Experimental results show that the hybrid LSTM-CNN-Attention model achieved outstanding performance, with an accuracy of 0.98, precision of 0.94, recall of 0.92, and F1-score of 0.93. These results surpass the performance of baseline models based solely on CNNs, LSTMs, or transformer-based classifiers such as BERT. The combination of neural network components enabled the model to effectively capture both fine-grained text structures and broader semantic context. Furthermore, the use of GloVe embeddings provided an efficient and effective representation of textual data, making the model suitable for integration into systems with real-time or near-real-time requirements. The proposed hybrid architecture demonstrates high effectiveness in text-based web content classification, particularly in tasks requiring both syntactic feature extraction and semantic interpretation. By combining presented mechanisms, the model addresses the limitations of individual architectures and achieves improved generalization. These findings support the broader use of hybrid deep learning approaches in NLP applications, especially where complex, unstructured textual data must be processed and classified with high reliability.
title Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
topic Computation and Language
Machine Learning
68T50 (Primary) 68T07 (Secondary)
I.2.7; I.2.6
url https://arxiv.org/abs/2512.18475