Saved in:
Bibliographic Details
Main Authors: Nauman, Mohd, Gvm, Sravan, Devane, Vijay, Pawar, Shyam, Thakur, Viraj, Pundalik, Kundeshwar, Sawarkar, Piyush, Saluja, Rohit, Desarkar, Maunendra, Ramakrishnan, Ganesh
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.02374
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912686945599488
author Nauman, Mohd
Gvm, Sravan
Devane, Vijay
Pawar, Shyam
Thakur, Viraj
Pundalik, Kundeshwar
Sawarkar, Piyush
Saluja, Rohit
Desarkar, Maunendra
Ramakrishnan, Ganesh
author_facet Nauman, Mohd
Gvm, Sravan
Devane, Vijay
Pawar, Shyam
Thakur, Viraj
Pundalik, Kundeshwar
Sawarkar, Piyush
Saluja, Rohit
Desarkar, Maunendra
Ramakrishnan, Ganesh
contents Current large language models excel at broad, general-purpose tasks, but consistently underperform when exposed to highly specialized domains that require deep cultural, linguistic, and subject-matter expertise. In particular, traditional medical systems such as Ayurveda embody centuries of nuanced textual and clinical knowledge that mainstream LLMs fail to accurately interpret or apply. We introduce AyurParam-2.9B, a domain-specialized, bilingual language model fine-tuned from Param-1-2.9B using an extensive, expertly curated Ayurveda dataset spanning classical texts and clinical guidance. AyurParam's dataset incorporates context-aware, reasoning, and objective-style Q&A in both English and Hindi, with rigorous annotation protocols for factual precision and instructional clarity. Benchmarked on BhashaBench-Ayur, AyurParam not only surpasses all open-source instruction-tuned models in its size class (1.5--3B parameters), but also demonstrates competitive or superior performance compared to much larger models. The results from AyurParam highlight the necessity for authentic domain adaptation and high-quality supervision in delivering reliable, culturally congruent AI for specialized medical knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
Nauman, Mohd
Gvm, Sravan
Devane, Vijay
Pawar, Shyam
Thakur, Viraj
Pundalik, Kundeshwar
Sawarkar, Piyush
Saluja, Rohit
Desarkar, Maunendra
Ramakrishnan, Ganesh
Computation and Language
Artificial Intelligence
Current large language models excel at broad, general-purpose tasks, but consistently underperform when exposed to highly specialized domains that require deep cultural, linguistic, and subject-matter expertise. In particular, traditional medical systems such as Ayurveda embody centuries of nuanced textual and clinical knowledge that mainstream LLMs fail to accurately interpret or apply. We introduce AyurParam-2.9B, a domain-specialized, bilingual language model fine-tuned from Param-1-2.9B using an extensive, expertly curated Ayurveda dataset spanning classical texts and clinical guidance. AyurParam's dataset incorporates context-aware, reasoning, and objective-style Q&A in both English and Hindi, with rigorous annotation protocols for factual precision and instructional clarity. Benchmarked on BhashaBench-Ayur, AyurParam not only surpasses all open-source instruction-tuned models in its size class (1.5--3B parameters), but also demonstrates competitive or superior performance compared to much larger models. The results from AyurParam highlight the necessity for authentic domain adaptation and high-quality supervision in delivering reliable, culturally congruent AI for specialized medical knowledge.
title AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.02374