Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2604.17134 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918518456320000 |
|---|---|
| author | Avram, Andrei-Marius Antonie, Aureliu Valentin Croitoru, Cosmin-Mircea Muntean, Vlad Andrei Cercel, Dumitru-Clementin |
| author_facet | Avram, Andrei-Marius Antonie, Aureliu Valentin Croitoru, Cosmin-Mircea Muntean, Vlad Andrei Cercel, Dumitru-Clementin |
| contents | We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews across three domains (books, movies, and music) and 202,141 unlabeled samples. To address cross-lingual and cross-domain challenges, we propose a multi-target adversarial training framework that employs loss reversal with meta-learned coefficients to dynamically balance sentiment discrimination with domain and language invariance. XLM-R achieves an F1-score of 66.23% with our approach, outperforming the baseline by 4.64%. Few-shot evaluation shows that Llama-3.1-8B achieves 58.43% F1-score, revealing a meaningful trade-off between the efficiency of prompting-based approaches and the higher performance of task-specific fine-tuning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_17134 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian Avram, Andrei-Marius Antonie, Aureliu Valentin Croitoru, Cosmin-Mircea Muntean, Vlad Andrei Cercel, Dumitru-Clementin Computation and Language We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews across three domains (books, movies, and music) and 202,141 unlabeled samples. To address cross-lingual and cross-domain challenges, we propose a multi-target adversarial training framework that employs loss reversal with meta-learned coefficients to dynamically balance sentiment discrimination with domain and language invariance. XLM-R achieves an F1-score of 66.23% with our approach, outperforming the baseline by 4.64%. Few-shot evaluation shows that Llama-3.1-8B achieves 58.43% F1-score, revealing a meaningful trade-off between the efficiency of prompting-based approaches and the higher performance of task-specific fine-tuning. |
| title | RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2604.17134 |