Hybrid Extractive Abstractive Summarization for Multilingual Sentiment Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Krasitskii, Mikhail, Sidorov, Grigori, Kolesnikova, Olga, Hernandez, Liliana Chanona, Gelbukh, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910994992726016
author Krasitskii, Mikhail
Sidorov, Grigori
Kolesnikova, Olga
Hernandez, Liliana Chanona
Gelbukh, Alexander
author_facet Krasitskii, Mikhail
Sidorov, Grigori
Kolesnikova, Olga
Hernandez, Liliana Chanona
Gelbukh, Alexander
contents We propose a hybrid approach for multilingual sentiment analysis that combines extractive and abstractive summarization to address the limitations of standalone methods. The model integrates TF-IDF-based extraction with a fine-tuned XLM-R abstractive module, enhanced by dynamic thresholding and cultural adaptation. Experiments across 10 languages show significant improvements over baselines, achieving 0.90 accuracy for English and 0.84 for low-resource languages. The approach also demonstrates 22% greater computational efficiency than traditional methods. Practical applications include real-time brand monitoring and cross-cultural discourse analysis. Future work will focus on optimization for low-resource languages via 8-bit quantization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06929
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hybrid Extractive Abstractive Summarization for Multilingual Sentiment Analysis
Krasitskii, Mikhail
Sidorov, Grigori
Kolesnikova, Olga
Hernandez, Liliana Chanona
Gelbukh, Alexander
Computation and Language
We propose a hybrid approach for multilingual sentiment analysis that combines extractive and abstractive summarization to address the limitations of standalone methods. The model integrates TF-IDF-based extraction with a fine-tuned XLM-R abstractive module, enhanced by dynamic thresholding and cultural adaptation. Experiments across 10 languages show significant improvements over baselines, achieving 0.90 accuracy for English and 0.84 for low-resource languages. The approach also demonstrates 22% greater computational efficiency than traditional methods. Practical applications include real-time brand monitoring and cross-cultural discourse analysis. Future work will focus on optimization for low-resource languages via 8-bit quantization.
title Hybrid Extractive Abstractive Summarization for Multilingual Sentiment Analysis
topic Computation and Language
url https://arxiv.org/abs/2506.06929