A Large-Scale Benchmark for Vietnamese Sentence Paraphrases

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nguyen, Sang Quang, Van Nguyen, Kiet
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917918574379008
author Nguyen, Sang Quang
Van Nguyen, Kiet
author_facet Nguyen, Sang Quang
Van Nguyen, Kiet
contents This paper presents ViSP, a high-quality Vietnamese dataset for sentence paraphrasing, consisting of 1.2M original-paraphrase pairs collected from various domains. The dataset was constructed using a hybrid approach that combines automatic paraphrase generation with manual evaluation to ensure high quality. We conducted experiments using methods such as back-translation, EDA, and baseline models like BART and T5, as well as large language models (LLMs), including GPT-4o, Gemini-1.5, Aya, Qwen-2.5, and Meta-Llama-3.1 variants. To the best of our knowledge, this is the first large-scale study on Vietnamese paraphrasing. We hope that our dataset and findings will serve as a valuable foundation for future research and applications in Vietnamese paraphrase tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Large-Scale Benchmark for Vietnamese Sentence Paraphrases
Nguyen, Sang Quang
Van Nguyen, Kiet
Computation and Language
This paper presents ViSP, a high-quality Vietnamese dataset for sentence paraphrasing, consisting of 1.2M original-paraphrase pairs collected from various domains. The dataset was constructed using a hybrid approach that combines automatic paraphrase generation with manual evaluation to ensure high quality. We conducted experiments using methods such as back-translation, EDA, and baseline models like BART and T5, as well as large language models (LLMs), including GPT-4o, Gemini-1.5, Aya, Qwen-2.5, and Meta-Llama-3.1 variants. To the best of our knowledge, this is the first large-scale study on Vietnamese paraphrasing. We hope that our dataset and findings will serve as a valuable foundation for future research and applications in Vietnamese paraphrase tasks.
title A Large-Scale Benchmark for Vietnamese Sentence Paraphrases
topic Computation and Language
url https://arxiv.org/abs/2502.07188