Sample-Efficient Language Model for Hinglish Conversational AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Sakshi, Prakash, Abhinav, Shah, Aakriti, Sachdeva, Chaitanya, Dumpala, Sanjana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912349342924800
author Singh, Sakshi
Prakash, Abhinav
Shah, Aakriti
Sachdeva, Chaitanya
Dumpala, Sanjana
author_facet Singh, Sakshi
Prakash, Abhinav
Shah, Aakriti
Sachdeva, Chaitanya
Dumpala, Sanjana
contents This paper presents our process for developing a sample-efficient language model for a conversational Hinglish chatbot. Hinglish, a code-mixed language that combines Hindi and English, presents a unique computational challenge due to inconsistent spelling, lack of standardization, and limited quality of conversational data. This work evaluates multiple pre-trained cross-lingual language models, including Gemma3-4B and Qwen2.5-7B, and employs fine-tuning techniques to improve performance on Hinglish conversational tasks. The proposed approach integrates synthetically generated dialogues with insights from existing Hinglish datasets to address data scarcity. Experimental results demonstrate that models with fewer parameters, when appropriately fine-tuned on high-quality code-mixed data, can achieve competitive performance for Hinglish conversation generation while maintaining computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sample-Efficient Language Model for Hinglish Conversational AI
Singh, Sakshi
Prakash, Abhinav
Shah, Aakriti
Sachdeva, Chaitanya
Dumpala, Sanjana
Computation and Language
I.2.7; I.2.6; H.5.2
This paper presents our process for developing a sample-efficient language model for a conversational Hinglish chatbot. Hinglish, a code-mixed language that combines Hindi and English, presents a unique computational challenge due to inconsistent spelling, lack of standardization, and limited quality of conversational data. This work evaluates multiple pre-trained cross-lingual language models, including Gemma3-4B and Qwen2.5-7B, and employs fine-tuning techniques to improve performance on Hinglish conversational tasks. The proposed approach integrates synthetically generated dialogues with insights from existing Hinglish datasets to address data scarcity. Experimental results demonstrate that models with fewer parameters, when appropriately fine-tuned on high-quality code-mixed data, can achieve competitive performance for Hinglish conversation generation while maintaining computational efficiency.
title Sample-Efficient Language Model for Hinglish Conversational AI
topic Computation and Language
I.2.7; I.2.6; H.5.2
url https://arxiv.org/abs/2504.19070