Multilingual Text Style Transfer: Datasets & Models for Indian Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mukherjee, Sourabrata, Ojha, Atul Kr., Bansal, Akanksha, Alok, Deepak, McCrae, John P., Dušek, Ondřej
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909296628858880
author Mukherjee, Sourabrata
Ojha, Atul Kr.
Bansal, Akanksha
Alok, Deepak
McCrae, John P.
Dušek, Ondřej
author_facet Mukherjee, Sourabrata
Ojha, Atul Kr.
Bansal, Akanksha
Alok, Deepak
McCrae, John P.
Dušek, Ondřej
contents Text style transfer (TST) involves altering the linguistic style of a text while preserving its core content. This paper focuses on sentiment transfer, a popular TST subtask, across a spectrum of Indian languages: Hindi, Magahi, Malayalam, Marathi, Punjabi, Odia, Telugu, and Urdu, expanding upon previous work on English-Bangla sentiment transfer (Mukherjee et al., 2023). We introduce dedicated datasets of 1,000 positive and 1,000 negative style-parallel sentences for each of these eight languages. We then evaluate the performance of various benchmark models categorized into parallel, non-parallel, cross-lingual, and shared learning approaches, including the Llama2 and GPT-3.5 large language models (LLMs). Our experiments highlight the significance of parallel data in TST and demonstrate the effectiveness of the Masked Style Filling (MSF) approach (Mukherjee et al., 2023) in non-parallel techniques. Moreover, cross-lingual and joint multilingual learning methods show promise, offering insights into selecting optimal models tailored to the specific language and task requirements. To the best of our knowledge, this work represents the first comprehensive exploration of the TST task as sentiment transfer across a diverse set of languages.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20805
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multilingual Text Style Transfer: Datasets & Models for Indian Languages
Mukherjee, Sourabrata
Ojha, Atul Kr.
Bansal, Akanksha
Alok, Deepak
McCrae, John P.
Dušek, Ondřej
Computation and Language
Text style transfer (TST) involves altering the linguistic style of a text while preserving its core content. This paper focuses on sentiment transfer, a popular TST subtask, across a spectrum of Indian languages: Hindi, Magahi, Malayalam, Marathi, Punjabi, Odia, Telugu, and Urdu, expanding upon previous work on English-Bangla sentiment transfer (Mukherjee et al., 2023). We introduce dedicated datasets of 1,000 positive and 1,000 negative style-parallel sentences for each of these eight languages. We then evaluate the performance of various benchmark models categorized into parallel, non-parallel, cross-lingual, and shared learning approaches, including the Llama2 and GPT-3.5 large language models (LLMs). Our experiments highlight the significance of parallel data in TST and demonstrate the effectiveness of the Masked Style Filling (MSF) approach (Mukherjee et al., 2023) in non-parallel techniques. Moreover, cross-lingual and joint multilingual learning methods show promise, offering insights into selecting optimal models tailored to the specific language and task requirements. To the best of our knowledge, this work represents the first comprehensive exploration of the TST task as sentiment transfer across a diverse set of languages.
title Multilingual Text Style Transfer: Datasets & Models for Indian Languages
topic Computation and Language
url https://arxiv.org/abs/2405.20805