Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sato, Soma, Tsukagoshi, Hayato, Sasano, Ryohei, Takeda, Koichi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929446258212864
author Sato, Soma
Tsukagoshi, Hayato
Sasano, Ryohei
Takeda, Koichi
author_facet Sato, Soma
Tsukagoshi, Hayato
Sasano, Ryohei
Takeda, Koichi
contents Decoder-based large language models (LLMs) have shown high performance on many tasks in natural language processing. This is also true for sentence embedding learning, where a decoder-based model, PromptEOL, has achieved the best performance on semantic textual similarity (STS) tasks. However, PromptEOL requires a manually annotated natural language inference (NLI) dataset for fine-tuning. We aim to improve sentence embeddings without using large manually annotated datasets by automatically generating an NLI dataset with an LLM and using it for fine-tuning of PromptEOL. To achieve this, we explore methods of data generation suitable for sentence embedding learning in this study. Specifically, we will focus on automatic dataset generation through few-shot learning and explore the appropriate methods to leverage few-shot examples. Experimental results on the STS tasks demonstrate that our approach outperforms existing models in settings without large manually annotated datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15132
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
Sato, Soma
Tsukagoshi, Hayato
Sasano, Ryohei
Takeda, Koichi
Computation and Language
Machine Learning
Decoder-based large language models (LLMs) have shown high performance on many tasks in natural language processing. This is also true for sentence embedding learning, where a decoder-based model, PromptEOL, has achieved the best performance on semantic textual similarity (STS) tasks. However, PromptEOL requires a manually annotated natural language inference (NLI) dataset for fine-tuning. We aim to improve sentence embeddings without using large manually annotated datasets by automatically generating an NLI dataset with an LLM and using it for fine-tuning of PromptEOL. To achieve this, we explore methods of data generation suitable for sentence embedding learning in this study. Specifically, we will focus on automatic dataset generation through few-shot learning and explore the appropriate methods to leverage few-shot examples. Experimental results on the STS tasks demonstrate that our approach outperforms existing models in settings without large manually annotated datasets.
title Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2402.15132