GFG -- Gender-Fair Generation: A CALAMITA Challenge

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Frenda, Simona, Piergentili, Andrea, Savoldi, Beatrice, Madeddu, Marco, Rosola, Martina, Casola, Silvia, Ferrando, Chiara, Patti, Viviana, Negri, Matteo, Bentivogli, Luisa
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910766091730944
author Frenda, Simona
Piergentili, Andrea
Savoldi, Beatrice
Madeddu, Marco
Rosola, Martina
Casola, Silvia
Ferrando, Chiara
Patti, Viviana
Negri, Matteo
Bentivogli, Luisa
author_facet Frenda, Simona
Piergentili, Andrea
Savoldi, Beatrice
Madeddu, Marco
Rosola, Martina
Casola, Silvia
Ferrando, Chiara
Patti, Viviana
Negri, Matteo
Bentivogli, Luisa
contents Gender-fair language aims at promoting gender equality by using terms and expressions that include all identities and avoid reinforcing gender stereotypes. Implementing gender-fair strategies is particularly challenging in heavily gender-marked languages, such as Italian. To address this, the Gender-Fair Generation challenge intends to help shift toward gender-fair language in written communication. The challenge, designed to assess and monitor the recognition and generation of gender-fair language in both mono- and cross-lingual scenarios, includes three tasks: (1) the detection of gendered expressions in Italian sentences, (2) the reformulation of gendered expressions into gender-fair alternatives, and (3) the generation of gender-fair language in automatic translation from English to Italian. The challenge relies on three different annotated datasets: the GFL-it corpus, which contains Italian texts extracted from administrative documents provided by the University of Brescia; GeNTE, a bilingual test set for gender-neutral rewriting and translation built upon a subset of the Europarl dataset; and Neo-GATE, a bilingual test set designed to assess the use of non-binary neomorphemes in Italian for both fair formulation and translation tasks. Finally, each task is evaluated with specific metrics: average of F1-score obtained by means of BERTScore computed on each entry of the datasets for task 1, an accuracy measured with a gender-neutral classifier, and a coverage-weighted accuracy for tasks 2 and 3.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GFG -- Gender-Fair Generation: A CALAMITA Challenge
Frenda, Simona
Piergentili, Andrea
Savoldi, Beatrice
Madeddu, Marco
Rosola, Martina
Casola, Silvia
Ferrando, Chiara
Patti, Viviana
Negri, Matteo
Bentivogli, Luisa
Computation and Language
Gender-fair language aims at promoting gender equality by using terms and expressions that include all identities and avoid reinforcing gender stereotypes. Implementing gender-fair strategies is particularly challenging in heavily gender-marked languages, such as Italian. To address this, the Gender-Fair Generation challenge intends to help shift toward gender-fair language in written communication. The challenge, designed to assess and monitor the recognition and generation of gender-fair language in both mono- and cross-lingual scenarios, includes three tasks: (1) the detection of gendered expressions in Italian sentences, (2) the reformulation of gendered expressions into gender-fair alternatives, and (3) the generation of gender-fair language in automatic translation from English to Italian. The challenge relies on three different annotated datasets: the GFL-it corpus, which contains Italian texts extracted from administrative documents provided by the University of Brescia; GeNTE, a bilingual test set for gender-neutral rewriting and translation built upon a subset of the Europarl dataset; and Neo-GATE, a bilingual test set designed to assess the use of non-binary neomorphemes in Italian for both fair formulation and translation tasks. Finally, each task is evaluated with specific metrics: average of F1-score obtained by means of BERTScore computed on each entry of the datasets for task 1, an accuracy measured with a gender-neutral classifier, and a coverage-weighted accuracy for tasks 2 and 3.
title GFG -- Gender-Fair Generation: A CALAMITA Challenge
topic Computation and Language
url https://arxiv.org/abs/2412.19168