Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
Fuente:
arXiv
Saved in:
| Main Authors: | Shrestha, Adarsha, Pokharel, Basanta, Shrestha, Binit, Adhikari, Smriti, Gothe, Dinesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NepaliGPT: A Generative Language Model for the Nepali Language
by: Pudasaini, Shushanta, et al.
Published: (2025)
by: Pudasaini, Shushanta, et al.
Published: (2025)
Whisper Finetuning on Nepali Language
by: Rijal, Sanjay, et al.
Published: (2024)
by: Rijal, Sanjay, et al.
Published: (2024)
Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali
by: Rimal, Ananda, et al.
Published: (2026)
by: Rimal, Ananda, et al.
Published: (2026)
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers
by: Dhakal, Prakash, et al.
Published: (2024)
by: Dhakal, Prakash, et al.
Published: (2024)
Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
by: Shrestha, Aayush M., et al.
Published: (2026)
by: Shrestha, Aayush M., et al.
Published: (2026)
MEME-Fusion@CHiPSAL 2026: Multimodal Ablation Study of Hate Detection and Sentiment Analysis on Nepali Memes
by: Wagle, Samir, et al.
Published: (2026)
by: Wagle, Samir, et al.
Published: (2026)
Mitigating Structural Noise in Low-Resource S2TT: An Optimized Cascaded Nepali-English Pipeline with Punctuation Restoration
by: Chongbang, Tangsang, et al.
Published: (2026)
by: Chongbang, Tangsang, et al.
Published: (2026)
Generative AI for Named Entity Recognition in Low-Resource Language Nepali
by: Neupane, Sameer, et al.
Published: (2025)
by: Neupane, Sameer, et al.
Published: (2025)
Advancing Voice Cloning for Nepali: Leveraging Transfer Learning in a Low-Resource Language
by: Karki, Manjil, et al.
Published: (2024)
by: Karki, Manjil, et al.
Published: (2024)
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
by: Ghimire, Rupak Raj, et al.
Published: (2026)
by: Ghimire, Rupak Raj, et al.
Published: (2026)
Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context
by: Pandey, Ashish, et al.
Published: (2026)
by: Pandey, Ashish, et al.
Published: (2026)
Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
by: Shrestha, Manil, et al.
Published: (2025)
by: Shrestha, Manil, et al.
Published: (2025)
Difficulty Estimation and Simplification of French Text Using LLMs
by: Jamet, Henri, et al.
Published: (2024)
by: Jamet, Henri, et al.
Published: (2024)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
by: Kim, Minwu, et al.
Published: (2025)
by: Kim, Minwu, et al.
Published: (2025)
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
by: Shrestha, Safal, et al.
Published: (2025)
by: Shrestha, Safal, et al.
Published: (2025)
Conformal Prediction for Risk-Controlled Medical Entity Extraction Across Clinical Domains
by: Shrestha, Manil, et al.
Published: (2026)
by: Shrestha, Manil, et al.
Published: (2026)
Batching BPE Tokenization Merges
by: Morgan, Alexander P.
Published: (2024)
by: Morgan, Alexander P.
Published: (2024)
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
by: Kim, Minwu, et al.
Published: (2026)
by: Kim, Minwu, et al.
Published: (2026)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
Prediction hubs are context-informed frequent tokens in LLMs
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
by: Witold, Waligóra
Published: (2024)
by: Witold, Waligóra
Published: (2024)
Jacobian Scopes: token-level causal attributions in LLMs
by: Liu, Toni J. B., et al.
Published: (2026)
by: Liu, Toni J. B., et al.
Published: (2026)
Creating and Evaluating Code-Mixed Nepali-English and Telugu-English Datasets for Abusive Language Detection Using Traditional and Deep Learning Models
by: Pandey, Manish, et al.
Published: (2025)
by: Pandey, Manish, et al.
Published: (2025)
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
by: Georgiou, Efthymios, et al.
Published: (2025)
by: Georgiou, Efthymios, et al.
Published: (2025)
ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
by: Liu, Chunhua, et al.
Published: (2025)
by: Liu, Chunhua, et al.
Published: (2025)
Evaluating Large Language Models' Responses to Sexual and Reproductive Health Queries in Nepali
by: Sharma, Medha, et al.
Published: (2026)
by: Sharma, Medha, et al.
Published: (2026)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
by: Asgari, Ehsaneddin, et al.
Published: (2025)
by: Asgari, Ehsaneddin, et al.
Published: (2025)
Nepali Sign Language Characters Recognition: Dataset Development and Deep Learning Approaches
by: Poudel, Birat, et al.
Published: (2025)
by: Poudel, Birat, et al.
Published: (2025)
A Survey on LLM-Assisted Clinical Trial Recruitment
by: Ghosh, Shrestha, et al.
Published: (2025)
by: Ghosh, Shrestha, et al.
Published: (2025)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
by: Shrestha, Safal, et al.
Published: (2025)
by: Shrestha, Safal, et al.
Published: (2025)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
by: Gautam, Aayush, et al.
Published: (2025)
by: Gautam, Aayush, et al.
Published: (2025)
Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search
by: Kim, Edward, et al.
Published: (2024)
by: Kim, Edward, et al.
Published: (2024)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
From Policy to Logic for Efficient and Interpretable Coverage Assessment
by: Pokharel, Rhitabrat, et al.
Published: (2026)
by: Pokharel, Rhitabrat, et al.
Published: (2026)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
by: Dhakal, Manish, et al.
Published: (2024)
by: Dhakal, Manish, et al.
Published: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution
by: Kwon, Deuksin, et al.
Published: (2026)
by: Kwon, Deuksin, et al.
Published: (2026)
BgGPT 1.0: Extending English-centric LLMs to other languages
by: Alexandrov, Anton, et al.
Published: (2024)
by: Alexandrov, Anton, et al.
Published: (2024)
Benchmarking GPT-5 for biomedical natural language processing
by: Hou, Yu, et al.
Published: (2025)
by: Hou, Yu, et al.
Published: (2025)
Similar Items
-
NepaliGPT: A Generative Language Model for the Nepali Language
by: Pudasaini, Shushanta, et al.
Published: (2025) -
Whisper Finetuning on Nepali Language
by: Rijal, Sanjay, et al.
Published: (2024) -
Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali
by: Rimal, Ananda, et al.
Published: (2026) -
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers
by: Dhakal, Prakash, et al.
Published: (2024) -
Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
by: Shrestha, Aayush M., et al.
Published: (2026)