Guardado en:
Detalles Bibliográficos
Autores principales: Avram, Andrei-Marius, Antonie, Aureliu Valentin, Croitoru, Cosmin-Mircea, Muntean, Vlad Andrei, Cercel, Dumitru-Clementin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2604.17134
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918518456320000
author Avram, Andrei-Marius
Antonie, Aureliu Valentin
Croitoru, Cosmin-Mircea
Muntean, Vlad Andrei
Cercel, Dumitru-Clementin
author_facet Avram, Andrei-Marius
Antonie, Aureliu Valentin
Croitoru, Cosmin-Mircea
Muntean, Vlad Andrei
Cercel, Dumitru-Clementin
contents We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews across three domains (books, movies, and music) and 202,141 unlabeled samples. To address cross-lingual and cross-domain challenges, we propose a multi-target adversarial training framework that employs loss reversal with meta-learned coefficients to dynamically balance sentiment discrimination with domain and language invariance. XLM-R achieves an F1-score of 66.23% with our approach, outperforming the baseline by 4.64%. Few-shot evaluation shows that Llama-3.1-8B achieves 58.43% F1-score, revealing a meaningful trade-off between the efficiency of prompting-based approaches and the higher performance of task-specific fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17134
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
Avram, Andrei-Marius
Antonie, Aureliu Valentin
Croitoru, Cosmin-Mircea
Muntean, Vlad Andrei
Cercel, Dumitru-Clementin
Computation and Language
We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews across three domains (books, movies, and music) and 202,141 unlabeled samples. To address cross-lingual and cross-domain challenges, we propose a multi-target adversarial training framework that employs loss reversal with meta-learned coefficients to dynamically balance sentiment discrimination with domain and language invariance. XLM-R achieves an F1-score of 66.23% with our approach, outperforming the baseline by 4.64%. Few-shot evaluation shows that Llama-3.1-8B achieves 58.43% F1-score, revealing a meaningful trade-off between the efficiency of prompting-based approaches and the higher performance of task-specific fine-tuning.
title RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
topic Computation and Language
url https://arxiv.org/abs/2604.17134