Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Christian, William, Adamlu, Daniel, Yu, Adrian, Suhartono, Derwin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914044585181184
author Christian, William
Adamlu, Daniel
Yu, Adrian
Suhartono, Derwin
author_facet Christian, William
Adamlu, Daniel
Yu, Adrian
Suhartono, Derwin
contents Understanding emotions in the Indonesian language is essential for improving customer experiences in e-commerce. This study focuses on enhancing the accuracy of emotion classification in Indonesian by leveraging advanced language models, IndoBERT and DistilBERT. A key component of our approach was data processing, specifically data augmentation, which included techniques such as back-translation and synonym replacement. These methods played a significant role in boosting the model's performance. After hyperparameter tuning, IndoBERT achieved an accuracy of 80\%, demonstrating the impact of careful data processing. While combining multiple IndoBERT models led to a slight improvement, it did not significantly enhance performance. Our findings indicate that IndoBERT was the most effective model for emotion classification in Indonesian, with data augmentation proving to be a vital factor in achieving high accuracy. Future research should focus on exploring alternative architectures and strategies to improve generalization for Indonesian NLP tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews
Christian, William
Adamlu, Daniel
Yu, Adrian
Suhartono, Derwin
Computation and Language
Understanding emotions in the Indonesian language is essential for improving customer experiences in e-commerce. This study focuses on enhancing the accuracy of emotion classification in Indonesian by leveraging advanced language models, IndoBERT and DistilBERT. A key component of our approach was data processing, specifically data augmentation, which included techniques such as back-translation and synonym replacement. These methods played a significant role in boosting the model's performance. After hyperparameter tuning, IndoBERT achieved an accuracy of 80\%, demonstrating the impact of careful data processing. While combining multiple IndoBERT models led to a slight improvement, it did not significantly enhance performance. Our findings indicate that IndoBERT was the most effective model for emotion classification in Indonesian, with data augmentation proving to be a vital factor in achieving high accuracy. Future research should focus on exploring alternative architectures and strategies to improve generalization for Indonesian NLP tasks.
title Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews
topic Computation and Language
url https://arxiv.org/abs/2509.14611