Low-Resource Transliteration for Roman-Urdu and Urdu Using Transformer-Based Models
Fuente:
arXiv
Saved in:
| Main Authors: | Butt, Umer, Veranasi, Stalin, Neumann, Günter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO
by: Butt, Umer, et al.
Published: (2024)
by: Butt, Umer, et al.
Published: (2024)
UrduLM: A Resource-Efficient Monolingual Urdu Language Model
by: Ali, Syed Muhammad, et al.
Published: (2026)
by: Ali, Syed Muhammad, et al.
Published: (2026)
AI-Generated Text Detection in Low-Resource Languages: A Case Study on Urdu
by: Ammar, Muhammad, et al.
Published: (2025)
by: Ammar, Muhammad, et al.
Published: (2025)
Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection
by: Harris, Sheetal, et al.
Published: (2024)
by: Harris, Sheetal, et al.
Published: (2024)
Hope Speech Detection in code-mixed Roman Urdu tweets: A Positive Turn in Natural Language Processing
by: Ahmad, Muhammad, et al.
Published: (2025)
by: Ahmad, Muhammad, et al.
Published: (2025)
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
by: Faisal, Faizan, et al.
Published: (2024)
by: Faisal, Faizan, et al.
Published: (2024)
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings
by: Fiaz, Layba, et al.
Published: (2025)
by: Fiaz, Layba, et al.
Published: (2025)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
by: Shafique, Muhammad Ali, et al.
Published: (2026)
by: Shafique, Muhammad Ali, et al.
Published: (2026)
UQA: Corpus for Urdu Question Answering
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
MUTEX: Leveraging Multilingual Transformers and Conditional Random Fields for Enhanced Urdu Toxic Span Detection
by: Arshad, Inayat, et al.
Published: (2026)
by: Arshad, Inayat, et al.
Published: (2026)
Enhanced Urdu Intent Detection with Large Language Models and Prototype-Informed Predictive Pipelines
by: Hassan, Faiza, et al.
Published: (2025)
by: Hassan, Faiza, et al.
Published: (2025)
Document-Level Sentiment Analysis of Urdu Text Using Deep Learning Techniques
by: Irum, Ammarah, et al.
Published: (2025)
by: Irum, Ammarah, et al.
Published: (2025)
ERUPD -- English to Roman Urdu Parallel Dataset
by: Furqan, Mohammed, et al.
Published: (2024)
by: Furqan, Mohammed, et al.
Published: (2024)
Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
by: Shafique, Muhammad Ali, et al.
Published: (2025)
by: Shafique, Muhammad Ali, et al.
Published: (2025)
Celebrity Profiling on Short Urdu Text using Twitter Followers' Feed
by: Hamza, Muhammad, et al.
Published: (2025)
by: Hamza, Muhammad, et al.
Published: (2025)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
A Literature Review of Keyword Spotting Technologies for Urdu
by: Rizvi, Syed Muhammad Aqdas
Published: (2024)
by: Rizvi, Syed Muhammad Aqdas
Published: (2024)
Assessing the Feasibility of Lightweight Whisper Models for Low-Resource Urdu Transcription
by: Antall, Abdul Rehman, et al.
Published: (2025)
by: Antall, Abdul Rehman, et al.
Published: (2025)
EDU-NER-2025: Named Entity Recognition in Urdu Educational Texts using XLM-RoBERTa with X (formerly Twitter)
by: Ullah, Fida, et al.
Published: (2025)
by: Ullah, Fida, et al.
Published: (2025)
Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training
by: Hassan, Muhammad Taimoor, et al.
Published: (2026)
by: Hassan, Muhammad Taimoor, et al.
Published: (2026)
AyutthayaAlpha: A Thai-Latin Script Transliteration Transformer
by: Lauc, Davor, et al.
Published: (2024)
by: Lauc, Davor, et al.
Published: (2024)
A Paradigm Gap in Urdu
by: Adeeba, Farah, et al.
Published: (2025)
by: Adeeba, Farah, et al.
Published: (2025)
ULTRA:Urdu Language Transformer-based Recommendation Architecture
by: Bashir, Alishbah, et al.
Published: (2026)
by: Bashir, Alishbah, et al.
Published: (2026)
Exploring the Role of Transliteration in In-Context Learning for Low-resource Languages Written in Non-Latin Scripts
by: Ma, Chunlan, et al.
Published: (2024)
by: Ma, Chunlan, et al.
Published: (2024)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025)
by: Ahmad, Sarfraz, et al.
Published: (2025)
Fake News Classification in Urdu: A Domain Adaptation Approach for a Low-Resource Language
by: Ali, Muhammad Zain, et al.
Published: (2025)
by: Ali, Muhammad Zain, et al.
Published: (2025)
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages
by: Azam, Gulfarogh, et al.
Published: (2025)
by: Azam, Gulfarogh, et al.
Published: (2025)
Evaluating Large Language Models on Urdu Idiom Translation
by: Khan, Muhammad Farmal, et al.
Published: (2025)
by: Khan, Muhammad Farmal, et al.
Published: (2025)
Happiness is Sharing a Vocabulary: A Study of Transliteration Methods
by: Jung, Haeji, et al.
Published: (2025)
by: Jung, Haeji, et al.
Published: (2025)
MALT: Mechanistic Ablation of Lossy Translation in LLMs for a Low-Resource Language: Urdu
by: Bajwa, Taaha Saleem
Published: (2025)
by: Bajwa, Taaha Saleem
Published: (2025)
NADIR: Differential Attention Flow for Non-Autoregressive Transliteration in Indic Languages
by: Tomar, Lakshya, et al.
Published: (2026)
by: Tomar, Lakshya, et al.
Published: (2026)
A Sentence‐Level Encoder–Decoder Architecture for Designing an Administrative Roman Urdu Chatbot
by: Muhammad Nazam Maqbool, et al.
Published: (2025)
by: Muhammad Nazam Maqbool, et al.
Published: (2025)
Transformers for Low-Resource Languages: Is Féidir Linn!
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
Generalists vs. Specialists: Evaluating Large Language Models for Urdu
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
by: Hassan, Umair
Published: (2025)
by: Hassan, Umair
Published: (2025)
A Permuted Autoregressive Approach to Word-Level Recognition for Urdu Digital Text
by: Mustafa, Ahmed, et al.
Published: (2024)
by: Mustafa, Ahmed, et al.
Published: (2024)
Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text
by: Hussain, Nisar, et al.
Published: (2025)
by: Hussain, Nisar, et al.
Published: (2025)
WER We Stand: Benchmarking Urdu ASR Models
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
by: Husain, Jaavid Aktar, et al.
Published: (2024)
by: Husain, Jaavid Aktar, et al.
Published: (2024)
Detection of Human and Machine-Authored Fake News in Urdu
by: Ali, Muhammad Zain, et al.
Published: (2024)
by: Ali, Muhammad Zain, et al.
Published: (2024)
Similar Items
-
Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO
by: Butt, Umer, et al.
Published: (2024) -
UrduLM: A Resource-Efficient Monolingual Urdu Language Model
by: Ali, Syed Muhammad, et al.
Published: (2026) -
AI-Generated Text Detection in Low-Resource Languages: A Case Study on Urdu
by: Ammar, Muhammad, et al.
Published: (2025) -
Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection
by: Harris, Sheetal, et al.
Published: (2024) -
Hope Speech Detection in code-mixed Roman Urdu tweets: A Positive Turn in Natural Language Processing
by: Ahmad, Muhammad, et al.
Published: (2025)