Development of Pre-Trained Transformer-based Models for the Nepali Language
Fuente:
arXiv
Saved in:
| Main Authors: | Thapa, Prajwal, Nyachhyon, Jinu, Sharma, Mridul, Bal, Bal Krishna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
by: Nyachhyon, Jinu, et al.
Published: (2024)
by: Nyachhyon, Jinu, et al.
Published: (2024)
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
by: Karki, Nischal, et al.
Published: (2026)
by: Karki, Nischal, et al.
Published: (2026)
Local Herb Identification Using Transfer Learning: A CNN-Powered Mobile Application for Nepalese Flora
by: Thapa, Prajwal, et al.
Published: (2025)
by: Thapa, Prajwal, et al.
Published: (2025)
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
by: Ghimire, Rupak Raj, et al.
Published: (2026)
by: Ghimire, Rupak Raj, et al.
Published: (2026)
Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
by: Begha, Funghang Limbu, et al.
Published: (2026)
by: Begha, Funghang Limbu, et al.
Published: (2026)
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
by: Ghimire, Rupak Raj, et al.
Published: (2024)
by: Ghimire, Rupak Raj, et al.
Published: (2024)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
NepaliGPT: A Generative Language Model for the Nepali Language
by: Pudasaini, Shushanta, et al.
Published: (2025)
by: Pudasaini, Shushanta, et al.
Published: (2025)
Can Perplexity Predict Fine-tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali
by: Luitel, Nishant, et al.
Published: (2024)
by: Luitel, Nishant, et al.
Published: (2024)
Machine Learning for Medical Billing Fraud and Insurance Risk Detection: Trends and Challenges in the US Healthcare System
by: Sharma, Dr. Bal Krishna
Published: (2025)
by: Sharma, Dr. Bal Krishna
Published: (2025)
GRASP: GRouped Activation Shared Parameterization for Parameter-Efficient Fine-Tuning and Robust Inference of Transformers
by: Bal, Malyaban, et al.
Published: (2025)
by: Bal, Malyaban, et al.
Published: (2025)
TharuChat: Bootstrapping Large Language Models for a Low-Resource Language via Synthetic Data and Human Validation
by: Panth, Prajwal, et al.
Published: (2026)
by: Panth, Prajwal, et al.
Published: (2026)
Domain-Adaptive Continued Pre-Training of Small Language Models
by: Faroz, Salman
Published: (2025)
by: Faroz, Salman
Published: (2025)
Domain-adaptative Continual Learning for Low-resource Tasks: Evaluation on Nepali
by: Duwal, Sharad, et al.
Published: (2024)
by: Duwal, Sharad, et al.
Published: (2024)
Cascade-Aware Training of Language Models
by: Wang, Congchao, et al.
Published: (2024)
by: Wang, Congchao, et al.
Published: (2024)
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
by: Sharma, Kartik, et al.
Published: (2025)
by: Sharma, Kartik, et al.
Published: (2025)
Safe Reinforcement Learning with Free-form Natural Language Constraints and Pre-Trained Language Models
by: Lou, Xingzhou, et al.
Published: (2024)
by: Lou, Xingzhou, et al.
Published: (2024)
Empirical Analysis of Efficient Fine-Tuning Methods for Large Pre-Trained Language Models
by: Doering, Nigel, et al.
Published: (2024)
by: Doering, Nigel, et al.
Published: (2024)
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
by: van der Wal, Oskar, et al.
Published: (2025)
by: van der Wal, Oskar, et al.
Published: (2025)
Automated Multi-Language to English Machine Translation Using Generative Pre-Trained Transformers
by: Pelofske, Elijah, et al.
Published: (2024)
by: Pelofske, Elijah, et al.
Published: (2024)
Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models
by: Zhang, Jingyang, et al.
Published: (2024)
by: Zhang, Jingyang, et al.
Published: (2024)
The Dark Side of the Language: Pre-trained Transformers in the DarkNet
by: Ranaldi, Leonardo, et al.
Published: (2022)
by: Ranaldi, Leonardo, et al.
Published: (2022)
Learning Dynamics in Continual Pre-Training for Large Language Models
by: Wang, Xingjin, et al.
Published: (2025)
by: Wang, Xingjin, et al.
Published: (2025)
Pre-Trained Policy Discriminators are General Reward Models
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
by: Sajith, Aryan, et al.
Published: (2024)
by: Sajith, Aryan, et al.
Published: (2024)
ForeCite: Adapting Pre-Trained Language Models to Predict Future Citation Rates of Academic Papers
by: Hull, Gavin, et al.
Published: (2025)
by: Hull, Gavin, et al.
Published: (2025)
Comparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition
by: Shirke, Mayur, et al.
Published: (2025)
by: Shirke, Mayur, et al.
Published: (2025)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
by: Schröder, Christopher, et al.
Published: (2024)
by: Schröder, Christopher, et al.
Published: (2024)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
NoteContrast: Contrastive Language-Diagnostic Pretraining for Medical Text
by: Kailas, Prajwal, et al.
Published: (2024)
by: Kailas, Prajwal, et al.
Published: (2024)
LLM-FS-Agent: A Deliberative Role-based Large Language Model Architecture for Transparent Feature Selection
by: Bal-Ghaoui, Mohamed, et al.
Published: (2025)
by: Bal-Ghaoui, Mohamed, et al.
Published: (2025)
TextAge: A Curated and Diverse Text Dataset for Age Classification
by: Cheekati, Shravan, et al.
Published: (2024)
by: Cheekati, Shravan, et al.
Published: (2024)
Model Merging in Pre-training of Large Language Models
by: Li, Yunshui, et al.
Published: (2025)
by: Li, Yunshui, et al.
Published: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
by: Lee, Tzu-Yun, et al.
Published: (2025)
by: Lee, Tzu-Yun, et al.
Published: (2025)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting
by: Cao, Defu, et al.
Published: (2023)
by: Cao, Defu, et al.
Published: (2023)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
by: Barron, Joshua, et al.
Published: (2025)
by: Barron, Joshua, et al.
Published: (2025)
Centered Masking for Language-Image Pre-Training
by: Liang, Mingliang, et al.
Published: (2024)
by: Liang, Mingliang, et al.
Published: (2024)
Small Vision-Language Models: A Survey on Compact Architectures and Techniques
by: Patnaik, Nitesh, et al.
Published: (2025)
by: Patnaik, Nitesh, et al.
Published: (2025)
Similar Items
-
Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
by: Nyachhyon, Jinu, et al.
Published: (2024) -
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
by: Karki, Nischal, et al.
Published: (2026) -
Local Herb Identification Using Transfer Learning: A CNN-Powered Mobile Application for Nepalese Flora
by: Thapa, Prajwal, et al.
Published: (2025) -
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
by: Ghimire, Rupak Raj, et al.
Published: (2026) -
Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
by: Begha, Funghang Limbu, et al.
Published: (2026)