IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Basak, Udvas, Dutta, Rajarshi, Pandey, Shivam, Modi, Ashutosh
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916196354359296
author Basak, Udvas
Dutta, Rajarshi
Pandey, Shivam
Modi, Ashutosh
author_facet Basak, Udvas
Dutta, Rajarshi
Pandey, Shivam
Modi, Ashutosh
contents This paper describes our system developed for the SemEval-2024 Task 1: Semantic Textual Relatedness. The challenge is focused on automatically detecting the degree of relatedness between pairs of sentences for 14 languages including both high and low-resource Asian and African languages. Our team participated in two subtasks consisting of Track A: supervised and Track B: unsupervised. This paper focuses on a BERT-based contrastive learning and similarity metric based approach primarily for the supervised track while exploring autoencoders for the unsupervised track. It also aims on the creation of a bigram relatedness corpus using negative sampling strategy, thereby producing refined word embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04513
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts
Basak, Udvas
Dutta, Rajarshi
Pandey, Shivam
Modi, Ashutosh
Computation and Language
Artificial Intelligence
Machine Learning
This paper describes our system developed for the SemEval-2024 Task 1: Semantic Textual Relatedness. The challenge is focused on automatically detecting the degree of relatedness between pairs of sentences for 14 languages including both high and low-resource Asian and African languages. Our team participated in two subtasks consisting of Track A: supervised and Track B: unsupervised. This paper focuses on a BERT-based contrastive learning and similarity metric based approach primarily for the supervised track while exploring autoencoders for the unsupervised track. It also aims on the creation of a bigram relatedness corpus using negative sampling strategy, thereby producing refined word embeddings.
title IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.04513