An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rusli, Andre, Shishido, Makoto
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929644694929408
author Rusli, Andre
Shishido, Makoto
author_facet Rusli, Andre
Shishido, Makoto
contents This study investigates the performance of three popular tokenization tools: MeCab, Sudachi, and SentencePiece, when applied as a preprocessing step for sentiment-based text classification of Japanese texts. Using Term Frequency-Inverse Document Frequency (TF-IDF) vectorization, we evaluate two traditional machine learning classifiers: Multinomial Naive Bayes and Logistic Regression. The results reveal that Sudachi produces tokens closely aligned with dictionary definitions, while MeCab and SentencePiece demonstrate faster processing speeds. The combination of SentencePiece, TF-IDF, and Logistic Regression outperforms the other alternatives in terms of classification performance.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17361
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification
Rusli, Andre
Shishido, Makoto
Computation and Language
This study investigates the performance of three popular tokenization tools: MeCab, Sudachi, and SentencePiece, when applied as a preprocessing step for sentiment-based text classification of Japanese texts. Using Term Frequency-Inverse Document Frequency (TF-IDF) vectorization, we evaluate two traditional machine learning classifiers: Multinomial Naive Bayes and Logistic Regression. The results reveal that Sudachi produces tokens closely aligned with dictionary definitions, while MeCab and SentencePiece demonstrate faster processing speeds. The combination of SentencePiece, TF-IDF, and Logistic Regression outperforms the other alternatives in terms of classification performance.
title An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification
topic Computation and Language
url https://arxiv.org/abs/2412.17361