PARAM-1 BharatGen 2.9B Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pundalik, Kundeshwar, Sawarkar, Piyush, Sahoo, Nihar, Shinde, Abhishek, Chanda, Prateek, Goswami, Vedant, Nagpal, Ajay, Singh, Atul, Thakur, Viraj, Dewane, Vijay, Thakur, Aamod, Patel, Bhargav, Gautam, Smita, Panditi, Bhagwan, Pawar, Shyam, Kotcha, Madhav, Racha, Suraj, Sureka, Saral, Singh, Pankaj, Bal, Rishi, Saluja, Rohit, Ramakrishnan, Ganesh
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916849683267584
author Pundalik, Kundeshwar
Sawarkar, Piyush
Sahoo, Nihar
Shinde, Abhishek
Chanda, Prateek
Goswami, Vedant
Nagpal, Ajay
Singh, Atul
Thakur, Viraj
Dewane, Vijay
Thakur, Aamod
Patel, Bhargav
Gautam, Smita
Panditi, Bhagwan
Pawar, Shyam
Kotcha, Madhav
Racha, Suraj
Sureka, Saral
Singh, Pankaj
Bal, Rishi
Saluja, Rohit
Ramakrishnan, Ganesh
author_facet Pundalik, Kundeshwar
Sawarkar, Piyush
Sahoo, Nihar
Shinde, Abhishek
Chanda, Prateek
Goswami, Vedant
Nagpal, Ajay
Singh, Atul
Thakur, Viraj
Dewane, Vijay
Thakur, Aamod
Patel, Bhargav
Gautam, Smita
Panditi, Bhagwan
Pawar, Shyam
Kotcha, Madhav
Racha, Suraj
Sureka, Saral
Singh, Pankaj
Bal, Rishi
Saluja, Rohit
Ramakrishnan, Ganesh
contents Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimization paradigms. This exclusionary design results in structural under-representation of linguistically diverse regions such as India, where over 20 official languages and 100+ dialects coexist alongside phenomena like code-switching and diglossia. We introduce PARAM-1, a 2.9B parameter decoder-only, text-only language model trained from scratch with an explicit architectural and linguistic focus on Indian diversity. PARAM-1 is trained on a bilingual dataset consisting of only Hindi and English, constructed with a strong focus on fact-rich, high-quality content. It is guided by three core principles: equitable representation of Indic languages through a 25% corpus allocation; tokenization fairness via a SentencePiece tokenizer adapted to Indian morphological structures; and culturally aligned evaluation benchmarks across IndicQA, code-mixed reasoning, and socio-linguistic robustness tasks. By embedding diversity at the pretraining level-rather than deferring it to post-hoc alignment-PARAM-1 offers a design-first blueprint for equitable foundation modeling. Our results demonstrate that it serves as both a competent general-purpose model and a robust baseline for India-centric applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PARAM-1 BharatGen 2.9B Model
Pundalik, Kundeshwar
Sawarkar, Piyush
Sahoo, Nihar
Shinde, Abhishek
Chanda, Prateek
Goswami, Vedant
Nagpal, Ajay
Singh, Atul
Thakur, Viraj
Dewane, Vijay
Thakur, Aamod
Patel, Bhargav
Gautam, Smita
Panditi, Bhagwan
Pawar, Shyam
Kotcha, Madhav
Racha, Suraj
Sureka, Saral
Singh, Pankaj
Bal, Rishi
Saluja, Rohit
Ramakrishnan, Ganesh
Computation and Language
Machine Learning
Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimization paradigms. This exclusionary design results in structural under-representation of linguistically diverse regions such as India, where over 20 official languages and 100+ dialects coexist alongside phenomena like code-switching and diglossia. We introduce PARAM-1, a 2.9B parameter decoder-only, text-only language model trained from scratch with an explicit architectural and linguistic focus on Indian diversity. PARAM-1 is trained on a bilingual dataset consisting of only Hindi and English, constructed with a strong focus on fact-rich, high-quality content. It is guided by three core principles: equitable representation of Indic languages through a 25% corpus allocation; tokenization fairness via a SentencePiece tokenizer adapted to Indian morphological structures; and culturally aligned evaluation benchmarks across IndicQA, code-mixed reasoning, and socio-linguistic robustness tasks. By embedding diversity at the pretraining level-rather than deferring it to post-hoc alignment-PARAM-1 offers a design-first blueprint for equitable foundation modeling. Our results demonstrate that it serves as both a competent general-purpose model and a robust baseline for India-centric applications.
title PARAM-1 BharatGen 2.9B Model
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.13390