PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
Fuente:
arXiv
Saved in:
| Main Author: | Rahman, Hanif |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seq vs Seq: An Open Suite of Paired Encoders and Decoders
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
by: Hasan, Khalid, et al.
Published: (2026)
by: Hasan, Khalid, et al.
Published: (2026)
PASC: Pipeline-Aware Conformal Prediction with Joint Coverage Guarantees for Multi-Stage NLP and LLM Pipelines
by: Kotte, Varun
Published: (2026)
by: Kotte, Varun
Published: (2026)
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
by: Ciancone, Mathieu, et al.
Published: (2024)
by: Ciancone, Mathieu, et al.
Published: (2024)
An Experimental Study on Data Augmentation Techniques for Named Entity Recognition on Low-Resource Domains
by: Torres, Arthur Elwing, et al.
Published: (2024)
by: Torres, Arthur Elwing, et al.
Published: (2024)
SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas
by: Wolff, Cornelius, et al.
Published: (2025)
by: Wolff, Cornelius, et al.
Published: (2025)
MoTime: A Dataset Suite for Multimodal Time Series Forecasting
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Twitter Sentiment Analysis using Distributed Word and Sentence Representation
by: Reddy, Dwarampudi Mahidhar, et al.
Published: (2019)
by: Reddy, Dwarampudi Mahidhar, et al.
Published: (2019)
One Word is Enough: Minimal Adversarial Perturbations for Neural Text Ranking
by: Karmakar, Tanmay, et al.
Published: (2026)
by: Karmakar, Tanmay, et al.
Published: (2026)
Urdu News Article Recommendation Model using Natural Language Processing Techniques
by: Abbas, Syed Zain, et al.
Published: (2022)
by: Abbas, Syed Zain, et al.
Published: (2022)
PluriHopRAG: Exhaustive, Recall-Sensitive QA Through Corpus-Specific Document Structure Learning
by: Sveistrys, Mykolas, et al.
Published: (2025)
by: Sveistrys, Mykolas, et al.
Published: (2025)
Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language
by: Rahman, Hanif, et al.
Published: (2026)
by: Rahman, Hanif, et al.
Published: (2026)
CAISSON: Concept-Augmented Inference Suite of Self-Organizing Neural Networks
by: Halperin, Igor
Published: (2024)
by: Halperin, Igor
Published: (2024)
UQA: Corpus for Urdu Question Answering
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
Information Extraction in Low-Resource Scenarios: Survey and Perspective
by: Deng, Shumin, et al.
Published: (2022)
by: Deng, Shumin, et al.
Published: (2022)
LegalRAG: A Hybrid RAG System for Multilingual Legal Information Retrieval
by: Kabir, Muhammad Rafsan, et al.
Published: (2025)
by: Kabir, Muhammad Rafsan, et al.
Published: (2025)
Reproducing HotFlip for Corpus Poisoning Attacks in Dense Retrieval
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes
by: Aqib, Mohammad, et al.
Published: (2025)
by: Aqib, Mohammad, et al.
Published: (2025)
IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
by: Paul, Shounak, et al.
Published: (2025)
by: Paul, Shounak, et al.
Published: (2025)
MAGE: Multi-Head Attention Guided Embeddings for Low Resource Sentiment Classification
by: Vashisht, Varun, et al.
Published: (2025)
by: Vashisht, Varun, et al.
Published: (2025)
FlashEvaluator: Expanding Search Space with Parallel Evaluation
by: Feng, Chao, et al.
Published: (2026)
by: Feng, Chao, et al.
Published: (2026)
AfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languages
by: Raju, Joshua Sakthivel, et al.
Published: (2025)
by: Raju, Joshua Sakthivel, et al.
Published: (2025)
Self-Augmented In-Context Learning for Unsupervised Word Translation
by: Li, Yaoyiran, et al.
Published: (2024)
by: Li, Yaoyiran, et al.
Published: (2024)
QueryBuilder: Human-in-the-Loop Query Development for Information Retrieval
by: Kandula, Hemanth, et al.
Published: (2024)
by: Kandula, Hemanth, et al.
Published: (2024)
ELMO: Efficiency via Low-precision and Peak Memory Optimization in Large Output Spaces
by: Zhang, Jinbin, et al.
Published: (2025)
by: Zhang, Jinbin, et al.
Published: (2025)
Improving Word Translation via Two-Stage Contrastive Learning
by: Li, Yaoyiran, et al.
Published: (2022)
by: Li, Yaoyiran, et al.
Published: (2022)
Language Models As Semantic Indexers
by: Jin, Bowen, et al.
Published: (2023)
by: Jin, Bowen, et al.
Published: (2023)
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
by: Weller, Orion, et al.
Published: (2024)
by: Weller, Orion, et al.
Published: (2024)
Evaluating Cost-Accuracy Trade-offs in Multimodal Search Relevance Judgements
by: Terragni, Silvia, et al.
Published: (2024)
by: Terragni, Silvia, et al.
Published: (2024)
From Topic to Transition Structure: Unsupervised Concept Discovery at Corpus Scale via Predictive Associative Memory
by: Dury, Jason
Published: (2026)
by: Dury, Jason
Published: (2026)
Make Large Language Model a Better Ranker
by: Chao, Wen-Shuo, et al.
Published: (2024)
by: Chao, Wen-Shuo, et al.
Published: (2024)
Accelerating Retrieval-Augmented Language Model Serving with Speculation
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
by: Zhang, Yunfan, et al.
Published: (2026)
by: Zhang, Yunfan, et al.
Published: (2026)
Evaluation of LLM-based Strategies for the Extraction of Food Product Information from Online Shops
by: Brosch, Christoph, et al.
Published: (2025)
by: Brosch, Christoph, et al.
Published: (2025)
Cascading Adaptors to Leverage English Data to Improve Performance of Question Answering for Low-Resource Languages
by: Pandya, Hariom A., et al.
Published: (2021)
by: Pandya, Hariom A., et al.
Published: (2021)
SLMRec: Distilling Large Language Models into Small for Sequential Recommendation
by: Xu, Wujiang, et al.
Published: (2024)
by: Xu, Wujiang, et al.
Published: (2024)
Optimizing Multi-Stage Language Models for Effective Text Retrieval
by: Trung, Quang Hoang, et al.
Published: (2024)
by: Trung, Quang Hoang, et al.
Published: (2024)
Integrating Large Language Models with Graphical Session-Based Recommendation
by: Guo, Naicheng, et al.
Published: (2024)
by: Guo, Naicheng, et al.
Published: (2024)
Can Small Language Models be Good Reasoners for Sequential Recommendation?
by: Wang, Yuling, et al.
Published: (2024)
by: Wang, Yuling, et al.
Published: (2024)
EvidenceRL: Reinforcing Evidence Consistency for Trustworthy Language Models
by: Tamo, J. Ben, et al.
Published: (2026)
by: Tamo, J. Ben, et al.
Published: (2026)
Similar Items
-
Seq vs Seq: An Open Suite of Paired Encoders and Decoders
by: Weller, Orion, et al.
Published: (2025) -
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
by: Hasan, Khalid, et al.
Published: (2026) -
PASC: Pipeline-Aware Conformal Prediction with Joint Coverage Guarantees for Multi-Stage NLP and LLM Pipelines
by: Kotte, Varun
Published: (2026) -
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
by: Ciancone, Mathieu, et al.
Published: (2024) -
An Experimental Study on Data Augmentation Techniques for Named Entity Recognition on Low-Resource Domains
by: Torres, Arthur Elwing, et al.
Published: (2024)