Foundational Study on Authorship Attribution of Japanese Web Reviews for Actor Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Matsubara, Hiroshi, Matsugaya, Shingo, Aoki, Taichi, Hashimoto, Masaki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914485018558464
author Matsubara, Hiroshi
Matsugaya, Shingo
Aoki, Taichi
Hashimoto, Masaki
author_facet Matsubara, Hiroshi
Matsugaya, Shingo
Aoki, Taichi
Hashimoto, Masaki
contents This study investigates the applicability of authorship attribution based on stylistic features to support actor analysis in threat intelligence. As a foundational step toward future application to dark web forums, we conducted experiments using Japanese review data from clear web sources. We constructed datasets from Rakuten Ichiba reviews and compared four methods: TF-IDF with logistic regression (TF-IDF+LR), BERT embeddings with logistic regression (BERT-Emb+LR), BERT fine-tuning (BERT-FT), and metric learning with $k$-nearest neighbors (Metric+kNN). Results showed that BERT-FT achieved the best performance; however, training became unstable as the number of authors scaled to several hundred, where TF-IDF+LR proved superior in terms of accuracy, stability, and computational cost. Furthermore, Top-$k$ evaluation demonstrated the utility of candidate screening, and error analysis revealed that boilerplate text, topic dependency, and short text length were primary factors causing misclassification.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16376
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Foundational Study on Authorship Attribution of Japanese Web Reviews for Actor Analysis
Matsubara, Hiroshi
Matsugaya, Shingo
Aoki, Taichi
Hashimoto, Masaki
Computation and Language
Cryptography and Security
This study investigates the applicability of authorship attribution based on stylistic features to support actor analysis in threat intelligence. As a foundational step toward future application to dark web forums, we conducted experiments using Japanese review data from clear web sources. We constructed datasets from Rakuten Ichiba reviews and compared four methods: TF-IDF with logistic regression (TF-IDF+LR), BERT embeddings with logistic regression (BERT-Emb+LR), BERT fine-tuning (BERT-FT), and metric learning with $k$-nearest neighbors (Metric+kNN). Results showed that BERT-FT achieved the best performance; however, training became unstable as the number of authors scaled to several hundred, where TF-IDF+LR proved superior in terms of accuracy, stability, and computational cost. Furthermore, Top-$k$ evaluation demonstrated the utility of candidate screening, and error analysis revealed that boilerplate text, topic dependency, and short text length were primary factors causing misclassification.
title Foundational Study on Authorship Attribution of Japanese Web Reviews for Actor Analysis
topic Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2604.16376