USTCCTSU at SemEval-2024 Task 1: Reducing Anisotropy for Cross-lingual Semantic Textual Relatedness Task

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jianjian, Liang, Shengwei, Liao, Yong, Deng, Hongping, Yu, Haiyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929608422588416
author Li, Jianjian
Liang, Shengwei
Liao, Yong
Deng, Hongping
Yu, Haiyang
author_facet Li, Jianjian
Liang, Shengwei
Liao, Yong
Deng, Hongping
Yu, Haiyang
contents Cross-lingual semantic textual relatedness task is an important research task that addresses challenges in cross-lingual communication and text understanding. It helps establish semantic connections between different languages, crucial for downstream tasks like machine translation, multilingual information retrieval, and cross-lingual text understanding.Based on extensive comparative experiments, we choose the XLM-R-base as our base model and use pre-trained sentence representations based on whitening to reduce anisotropy.Additionally, for the given training data, we design a delicate data filtering method to alleviate the curse of multilingualism. With our approach, we achieve a 2nd score in Spanish, a 3rd in Indonesian, and multiple entries in the top ten results in the competition's track C. We further do a comprehensive analysis to inspire future research aimed at improving performance on cross-lingual tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18990
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle USTCCTSU at SemEval-2024 Task 1: Reducing Anisotropy for Cross-lingual Semantic Textual Relatedness Task
Li, Jianjian
Liang, Shengwei
Liao, Yong
Deng, Hongping
Yu, Haiyang
Computation and Language
Artificial Intelligence
I.2.7
Cross-lingual semantic textual relatedness task is an important research task that addresses challenges in cross-lingual communication and text understanding. It helps establish semantic connections between different languages, crucial for downstream tasks like machine translation, multilingual information retrieval, and cross-lingual text understanding.Based on extensive comparative experiments, we choose the XLM-R-base as our base model and use pre-trained sentence representations based on whitening to reduce anisotropy.Additionally, for the given training data, we design a delicate data filtering method to alleviate the curse of multilingualism. With our approach, we achieve a 2nd score in Spanish, a 3rd in Indonesian, and multiple entries in the top ten results in the competition's track C. We further do a comprehensive analysis to inspire future research aimed at improving performance on cross-lingual tasks.
title USTCCTSU at SemEval-2024 Task 1: Reducing Anisotropy for Cross-lingual Semantic Textual Relatedness Task
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2411.18990