Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Nyachhyon, Jinu, Sharma, Mridul, Thapa, Prajwal, Bal, Bal Krishna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Development of Pre-Trained Transformer-based Models for the Nepali Language
by: Thapa, Prajwal, et al.
Published: (2024)
by: Thapa, Prajwal, et al.
Published: (2024)
Local Herb Identification Using Transfer Learning: A CNN-Powered Mobile Application for Nepalese Flora
by: Thapa, Prajwal, et al.
Published: (2025)
by: Thapa, Prajwal, et al.
Published: (2025)
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
by: Karki, Nischal, et al.
Published: (2026)
by: Karki, Nischal, et al.
Published: (2026)
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
by: Ghimire, Rupak Raj, et al.
Published: (2024)
by: Ghimire, Rupak Raj, et al.
Published: (2024)
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
by: Ghimire, Rupak Raj, et al.
Published: (2026)
by: Ghimire, Rupak Raj, et al.
Published: (2026)
Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
by: Begha, Funghang Limbu, et al.
Published: (2026)
by: Begha, Funghang Limbu, et al.
Published: (2026)
NepaliGPT: A Generative Language Model for the Nepali Language
by: Pudasaini, Shushanta, et al.
Published: (2025)
by: Pudasaini, Shushanta, et al.
Published: (2025)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
SymBa: Symbolic Backward Chaining for Structured Natural Language Reasoning
by: Lee, Jinu, et al.
Published: (2024)
by: Lee, Jinu, et al.
Published: (2024)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Benchmarking Retrieval Strategies for Biomedical Retrieval-Augmented Generation: A Controlled Empirical Study
by: Bal, Devi Prasad, et al.
Published: (2026)
by: Bal, Devi Prasad, et al.
Published: (2026)
Whisper Finetuning on Nepali Language
by: Rijal, Sanjay, et al.
Published: (2024)
by: Rijal, Sanjay, et al.
Published: (2024)
Confirmation bias: A challenge for scalable oversight
by: Recchia, Gabriel, et al.
Published: (2025)
by: Recchia, Gabriel, et al.
Published: (2025)
Evaluating Large Language Models' Responses to Sexual and Reproductive Health Queries in Nepali
by: Sharma, Medha, et al.
Published: (2026)
by: Sharma, Medha, et al.
Published: (2026)
Machine Learning for Medical Billing Fraud and Insurance Risk Detection: Trends and Challenges in the US Healthcare System
by: Sharma, Dr. Bal Krishna
Published: (2025)
by: Sharma, Dr. Bal Krishna
Published: (2025)
Domain-adaptative Continual Learning for Low-resource Tasks: Evaluation on Nepali
by: Duwal, Sharad, et al.
Published: (2024)
by: Duwal, Sharad, et al.
Published: (2024)
Hybrid EEG--Driven Brain--Computer Interface: A Large Language Model Framework for Personalized Language Rehabilitation
by: Hossain, Ismail, et al.
Published: (2025)
by: Hossain, Ismail, et al.
Published: (2025)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives
by: Su, Kun, et al.
Published: (2025)
by: Su, Kun, et al.
Published: (2025)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
by: Arora, Siddhant, et al.
Published: (2023)
by: Arora, Siddhant, et al.
Published: (2023)
Entailment-Preserving First-order Logic Representations in Natural Language Entailment
by: Lee, Jinu, et al.
Published: (2025)
by: Lee, Jinu, et al.
Published: (2025)
P-SpikeSSM: Harnessing Probabilistic Spiking State Space Models for Long-Range Dependency Tasks
by: Bal, Malyaban, et al.
Published: (2024)
by: Bal, Malyaban, et al.
Published: (2024)
Generative AI for Named Entity Recognition in Low-Resource Language Nepali
by: Neupane, Sameer, et al.
Published: (2025)
by: Neupane, Sameer, et al.
Published: (2025)
ArabicNLU 2024: The First Arabic Natural Language Understanding Shared Task
by: Khalilia, Mohammed, et al.
Published: (2024)
by: Khalilia, Mohammed, et al.
Published: (2024)
Nepali Sign Language Characters Recognition: Dataset Development and Deep Learning Approaches
by: Poudel, Birat, et al.
Published: (2025)
by: Poudel, Birat, et al.
Published: (2025)
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
by: Nonesung, Surapon, et al.
Published: (2025)
by: Nonesung, Surapon, et al.
Published: (2025)
Evaluating Step-by-step Reasoning Traces: A Survey
by: Lee, Jinu, et al.
Published: (2025)
by: Lee, Jinu, et al.
Published: (2025)
Swift Cross-Dataset Pruning: Enhancing Fine-Tuning Efficiency in Natural Language Understanding
by: Nguyen, Binh-Nguyen, et al.
Published: (2025)
by: Nguyen, Binh-Nguyen, et al.
Published: (2025)
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
by: Wang, Yubo, et al.
Published: (2024)
by: Wang, Yubo, et al.
Published: (2024)
A New Benchmark Dataset and Mixture-of-Experts Language Models for Adversarial Natural Language Inference in Vietnamese
by: Van Huynh, Tin, et al.
Published: (2024)
by: Van Huynh, Tin, et al.
Published: (2024)
New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark
by: Alavoine, Nadège, et al.
Published: (2024)
by: Alavoine, Nadège, et al.
Published: (2024)
TextAge: A Curated and Diverse Text Dataset for Age Classification
by: Cheekati, Shravan, et al.
Published: (2024)
by: Cheekati, Shravan, et al.
Published: (2024)
Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
by: Sharma, Swati, et al.
Published: (2026)
by: Sharma, Swati, et al.
Published: (2026)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
by: De, Anik, et al.
Published: (2025)
by: De, Anik, et al.
Published: (2025)
Optical Text Recognition in Nepali and Bengali: A Transformer-based Approach
by: Hasan, S M Rakib, et al.
Published: (2024)
by: Hasan, S M Rakib, et al.
Published: (2024)
SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation
by: Bal, Malyaban, et al.
Published: (2023)
by: Bal, Malyaban, et al.
Published: (2023)
Understanding the Dataset Practitioners Behind Large Language Model Development
by: Qian, Crystal, et al.
Published: (2024)
by: Qian, Crystal, et al.
Published: (2024)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models
by: Racha, Suraj, et al.
Published: (2025)
by: Racha, Suraj, et al.
Published: (2025)
INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects
by: Sharma, Tarun, et al.
Published: (2026)
by: Sharma, Tarun, et al.
Published: (2026)
Similar Items
-
Development of Pre-Trained Transformer-based Models for the Nepali Language
by: Thapa, Prajwal, et al.
Published: (2024) -
Local Herb Identification Using Transfer Learning: A CNN-Powered Mobile Application for Nepalese Flora
by: Thapa, Prajwal, et al.
Published: (2025) -
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
by: Karki, Nischal, et al.
Published: (2026) -
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
by: Ghimire, Rupak Raj, et al.
Published: (2024) -
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
by: Ghimire, Rupak Raj, et al.
Published: (2026)