PARAM-1 BharatGen 2.9B Model
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916849683267584 |
|---|---|
| author | Pundalik, Kundeshwar Sawarkar, Piyush Sahoo, Nihar Shinde, Abhishek Chanda, Prateek Goswami, Vedant Nagpal, Ajay Singh, Atul Thakur, Viraj Dewane, Vijay Thakur, Aamod Patel, Bhargav Gautam, Smita Panditi, Bhagwan Pawar, Shyam Kotcha, Madhav Racha, Suraj Sureka, Saral Singh, Pankaj Bal, Rishi Saluja, Rohit Ramakrishnan, Ganesh |
| author_facet | Pundalik, Kundeshwar Sawarkar, Piyush Sahoo, Nihar Shinde, Abhishek Chanda, Prateek Goswami, Vedant Nagpal, Ajay Singh, Atul Thakur, Viraj Dewane, Vijay Thakur, Aamod Patel, Bhargav Gautam, Smita Panditi, Bhagwan Pawar, Shyam Kotcha, Madhav Racha, Suraj Sureka, Saral Singh, Pankaj Bal, Rishi Saluja, Rohit Ramakrishnan, Ganesh |
| contents | Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimization paradigms. This exclusionary design results in structural under-representation of linguistically diverse regions such as India, where over 20 official languages and 100+ dialects coexist alongside phenomena like code-switching and diglossia. We introduce PARAM-1, a 2.9B parameter decoder-only, text-only language model trained from scratch with an explicit architectural and linguistic focus on Indian diversity. PARAM-1 is trained on a bilingual dataset consisting of only Hindi and English, constructed with a strong focus on fact-rich, high-quality content. It is guided by three core principles: equitable representation of Indic languages through a 25% corpus allocation; tokenization fairness via a SentencePiece tokenizer adapted to Indian morphological structures; and culturally aligned evaluation benchmarks across IndicQA, code-mixed reasoning, and socio-linguistic robustness tasks. By embedding diversity at the pretraining level-rather than deferring it to post-hoc alignment-PARAM-1 offers a design-first blueprint for equitable foundation modeling. Our results demonstrate that it serves as both a competent general-purpose model and a robust baseline for India-centric applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_13390 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | PARAM-1 BharatGen 2.9B Model Pundalik, Kundeshwar Sawarkar, Piyush Sahoo, Nihar Shinde, Abhishek Chanda, Prateek Goswami, Vedant Nagpal, Ajay Singh, Atul Thakur, Viraj Dewane, Vijay Thakur, Aamod Patel, Bhargav Gautam, Smita Panditi, Bhagwan Pawar, Shyam Kotcha, Madhav Racha, Suraj Sureka, Saral Singh, Pankaj Bal, Rishi Saluja, Rohit Ramakrishnan, Ganesh Computation and Language Machine Learning Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimization paradigms. This exclusionary design results in structural under-representation of linguistically diverse regions such as India, where over 20 official languages and 100+ dialects coexist alongside phenomena like code-switching and diglossia. We introduce PARAM-1, a 2.9B parameter decoder-only, text-only language model trained from scratch with an explicit architectural and linguistic focus on Indian diversity. PARAM-1 is trained on a bilingual dataset consisting of only Hindi and English, constructed with a strong focus on fact-rich, high-quality content. It is guided by three core principles: equitable representation of Indic languages through a 25% corpus allocation; tokenization fairness via a SentencePiece tokenizer adapted to Indian morphological structures; and culturally aligned evaluation benchmarks across IndicQA, code-mixed reasoning, and socio-linguistic robustness tasks. By embedding diversity at the pretraining level-rather than deferring it to post-hoc alignment-PARAM-1 offers a design-first blueprint for equitable foundation modeling. Our results demonstrate that it serves as both a competent general-purpose model and a robust baseline for India-centric applications. |
| title | PARAM-1 BharatGen 2.9B Model |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2507.13390 |