DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Lung-Hao, Yu, Liang-Chih, Loukashevich, Natalia, Alimova, Ilseyar, Panchenko, Alexander, Lin, Tzu-Mi, Xu, Zhe-Yu, Zhou, Jian-Yu, Zheng, Guangmin, Wang, Jin, Awasthi, Sharanya, Becker, Jonas, Wahle, Jan Philip, Ruas, Terry, Muhammad, Shamsuddeen Hassan, Mohammad, Saif M.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910161316085760
author Lee, Lung-Hao
Yu, Liang-Chih
Loukashevich, Natalia
Alimova, Ilseyar
Panchenko, Alexander
Lin, Tzu-Mi
Xu, Zhe-Yu
Zhou, Jian-Yu
Zheng, Guangmin
Wang, Jin
Awasthi, Sharanya
Becker, Jonas
Wahle, Jan Philip
Ruas, Terry
Muhammad, Shamsuddeen Hassan
Mohammad, Saif M.
author_facet Lee, Lung-Hao
Yu, Liang-Chih
Loukashevich, Natalia
Alimova, Ilseyar
Panchenko, Alexander
Lin, Tzu-Mi
Xu, Zhe-Yu
Zhou, Jian-Yu
Zheng, Guangmin
Wang, Jin
Awasthi, Sharanya
Becker, Jonas
Wahle, Jan Philip
Ruas, Terry
Muhammad, Shamsuddeen Hassan
Mohammad, Saif M.
contents Aspect-Based Sentiment Analysis (ABSA) focuses on extracting sentiment at a fine-grained aspect level and has been widely applied across real-world domains. However, existing ABSA research relies on coarse-grained categorical labels (e.g., positive, negative), which limits its ability to capture nuanced affective states. To address this limitation, we adopt a dimensional approach that represents sentiment with continuous valence-arousal (VA) scores, enabling fine-grained analysis at both the aspect and sentiment levels. To this end, we introduce DimABSA, the first multilingual, dimensional ABSA resource annotated with both traditional ABSA elements (aspect terms, aspect categories, and opinion terms) and newly introduced VA scores. This resource contains 76,958 aspect instances across 42,590 sentences, spanning six languages and four domains. We further introduce three subtasks that combine VA scores with different ABSA elements, providing a bridge from traditional ABSA to dimensional ABSA. Given that these subtasks involve both categorical and continuous outputs, we propose a new unified metric, continuous F1 (cF1), which incorporates VA prediction error into standard F1. We provide a comprehensive benchmark using both prompted and fine-tuned large language models across all subtasks. Our results show that DimABSA is a challenging benchmark and provides a foundation for advancing multilingual dimensional ABSA. We publicly released the DimABSA dataset, which was used for Track A of SemEval-2026 Task 3, attracting over 300 participants.
format Preprint
id arxiv_https___arxiv_org_abs_2601_23022
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
Lee, Lung-Hao
Yu, Liang-Chih
Loukashevich, Natalia
Alimova, Ilseyar
Panchenko, Alexander
Lin, Tzu-Mi
Xu, Zhe-Yu
Zhou, Jian-Yu
Zheng, Guangmin
Wang, Jin
Awasthi, Sharanya
Becker, Jonas
Wahle, Jan Philip
Ruas, Terry
Muhammad, Shamsuddeen Hassan
Mohammad, Saif M.
Computation and Language
I.2.7
Aspect-Based Sentiment Analysis (ABSA) focuses on extracting sentiment at a fine-grained aspect level and has been widely applied across real-world domains. However, existing ABSA research relies on coarse-grained categorical labels (e.g., positive, negative), which limits its ability to capture nuanced affective states. To address this limitation, we adopt a dimensional approach that represents sentiment with continuous valence-arousal (VA) scores, enabling fine-grained analysis at both the aspect and sentiment levels. To this end, we introduce DimABSA, the first multilingual, dimensional ABSA resource annotated with both traditional ABSA elements (aspect terms, aspect categories, and opinion terms) and newly introduced VA scores. This resource contains 76,958 aspect instances across 42,590 sentences, spanning six languages and four domains. We further introduce three subtasks that combine VA scores with different ABSA elements, providing a bridge from traditional ABSA to dimensional ABSA. Given that these subtasks involve both categorical and continuous outputs, we propose a new unified metric, continuous F1 (cF1), which incorporates VA prediction error into standard F1. We provide a comprehensive benchmark using both prompted and fine-tuned large language models across all subtasks. Our results show that DimABSA is a challenging benchmark and provides a foundation for advancing multilingual dimensional ABSA. We publicly released the DimABSA dataset, which was used for Track A of SemEval-2026 Task 3, attracting over 300 participants.
title DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2601.23022