Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kaur, Rishemjit, Bhankhar, Arshdeep Singh, Salh, Jashanpreet Singh, Rajput, Sudhir, Vidhi, Mahendra, Kashish, Berwal, Bhavika, Kumar, Ritesh, Ranathunga, Surangika
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915421003710464
author Kaur, Rishemjit
Bhankhar, Arshdeep Singh
Salh, Jashanpreet Singh
Rajput, Sudhir
Vidhi
Mahendra, Kashish
Berwal, Bhavika
Kumar, Ritesh
Ranathunga, Surangika
author_facet Kaur, Rishemjit
Bhankhar, Arshdeep Singh
Salh, Jashanpreet Singh
Rajput, Sudhir
Vidhi
Mahendra, Kashish
Berwal, Bhavika
Kumar, Ritesh
Ranathunga, Surangika
contents Enabling farmers to access accurate agriculture-related information in their native languages in a timely manner is crucial for the success of the agriculture field. Publicly available general-purpose Large Language Models (LLMs) typically offer generic agriculture advisories, lacking precision in local and multilingual contexts. Our study addresses this limitation by generating multilingual (English, Hindi, Punjabi) synthetic datasets from agriculture-specific documents from India and fine-tuning LLMs for the task of question answering (QA). Evaluation on human-created datasets demonstrates significant improvements in factuality, relevance, and agricultural consensus for the fine-tuned LLMs compared to the baseline counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16974
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain
Kaur, Rishemjit
Bhankhar, Arshdeep Singh
Salh, Jashanpreet Singh
Rajput, Sudhir
Vidhi
Mahendra, Kashish
Berwal, Bhavika
Kumar, Ritesh
Ranathunga, Surangika
Computation and Language
Artificial Intelligence
I.2.7; J.m
Enabling farmers to access accurate agriculture-related information in their native languages in a timely manner is crucial for the success of the agriculture field. Publicly available general-purpose Large Language Models (LLMs) typically offer generic agriculture advisories, lacking precision in local and multilingual contexts. Our study addresses this limitation by generating multilingual (English, Hindi, Punjabi) synthetic datasets from agriculture-specific documents from India and fine-tuning LLMs for the task of question answering (QA). Evaluation on human-created datasets demonstrates significant improvements in factuality, relevance, and agricultural consensus for the fine-tuned LLMs compared to the baseline counterparts.
title Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain
topic Computation and Language
Artificial Intelligence
I.2.7; J.m
url https://arxiv.org/abs/2507.16974