Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929446258212864 |
|---|---|
| author | Sato, Soma Tsukagoshi, Hayato Sasano, Ryohei Takeda, Koichi |
| author_facet | Sato, Soma Tsukagoshi, Hayato Sasano, Ryohei Takeda, Koichi |
| contents | Decoder-based large language models (LLMs) have shown high performance on many tasks in natural language processing. This is also true for sentence embedding learning, where a decoder-based model, PromptEOL, has achieved the best performance on semantic textual similarity (STS) tasks. However, PromptEOL requires a manually annotated natural language inference (NLI) dataset for fine-tuning. We aim to improve sentence embeddings without using large manually annotated datasets by automatically generating an NLI dataset with an LLM and using it for fine-tuning of PromptEOL. To achieve this, we explore methods of data generation suitable for sentence embedding learning in this study. Specifically, we will focus on automatic dataset generation through few-shot learning and explore the appropriate methods to leverage few-shot examples. Experimental results on the STS tasks demonstrate that our approach outperforms existing models in settings without large manually annotated datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_15132 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples Sato, Soma Tsukagoshi, Hayato Sasano, Ryohei Takeda, Koichi Computation and Language Machine Learning Decoder-based large language models (LLMs) have shown high performance on many tasks in natural language processing. This is also true for sentence embedding learning, where a decoder-based model, PromptEOL, has achieved the best performance on semantic textual similarity (STS) tasks. However, PromptEOL requires a manually annotated natural language inference (NLI) dataset for fine-tuning. We aim to improve sentence embeddings without using large manually annotated datasets by automatically generating an NLI dataset with an LLM and using it for fine-tuning of PromptEOL. To achieve this, we explore methods of data generation suitable for sentence embedding learning in this study. Specifically, we will focus on automatic dataset generation through few-shot learning and explore the appropriate methods to leverage few-shot examples. Experimental results on the STS tasks demonstrate that our approach outperforms existing models in settings without large manually annotated datasets. |
| title | Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2402.15132 |