Improving Informally Romanized Language Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Benton, Adrian, Gutkin, Alexander, Kirov, Christo, Roark, Brian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Graphemic Normalization of the Perso-Arabic Script
by: Doctor, Raiomond, et al.
Published: (2022)
by: Doctor, Raiomond, et al.
Published: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023)
by: Gutkin, Alexander, et al.
Published: (2023)
Paradigm Completion for Derivational Morphology
by: Cotterell, Ryan, et al.
Published: (2017)
by: Cotterell, Ryan, et al.
Published: (2017)
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages
by: Ruder, Sebastian, et al.
Published: (2023)
by: Ruder, Sebastian, et al.
Published: (2023)
Exploring and Improving Drafts in Blockwise Parallel Decoding
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
Toward Informal Language Processing: Knowledge of Slang in Large Language Models
by: Sun, Zhewei, et al.
Published: (2024)
by: Sun, Zhewei, et al.
Published: (2024)
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
by: Husain, Jaavid Aktar, et al.
Published: (2024)
by: Husain, Jaavid Aktar, et al.
Published: (2024)
One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models
by: Ebing, Benedikt, et al.
Published: (2026)
by: Ebing, Benedikt, et al.
Published: (2026)
A Universal Vibe? Finding and Controlling Language-Agnostic Informal Register with SAEs
by: Kialy, Uri Z., et al.
Published: (2026)
by: Kialy, Uri Z., et al.
Published: (2026)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
by: Cao, Jialun, et al.
Published: (2025)
by: Cao, Jialun, et al.
Published: (2025)
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report
by: Fedorova, Mariia, et al.
Published: (2026)
by: Fedorova, Mariia, et al.
Published: (2026)
GIFT: Games as Informal Training for Generalizable LLMs
by: Lyu, Nuoyan, et al.
Published: (2026)
by: Lyu, Nuoyan, et al.
Published: (2026)
Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
by: Christian, William, et al.
Published: (2025)
by: Christian, William, et al.
Published: (2025)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
by: Jia, Jinghan, et al.
Published: (2026)
by: Jia, Jinghan, et al.
Published: (2026)
Bilingual Word Level Language Identification for Omotic Languages
by: Yigezu, Mesay Gemeda, et al.
Published: (2025)
by: Yigezu, Mesay Gemeda, et al.
Published: (2025)
Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text
by: Hussain, Nisar, et al.
Published: (2025)
by: Hussain, Nisar, et al.
Published: (2025)
ERUPD -- English to Roman Urdu Parallel Dataset
by: Furqan, Mohammed, et al.
Published: (2024)
by: Furqan, Mohammed, et al.
Published: (2024)
Natural Language Translation of Formal Proofs through Informalization of Proof Steps and Recursive Summarization along Proof Structure
by: Hattori, Seiji, et al.
Published: (2025)
by: Hattori, Seiji, et al.
Published: (2025)
Advancing Bangla Machine Translation Through Informal Datasets
by: Roy, Ayon, et al.
Published: (2025)
by: Roy, Ayon, et al.
Published: (2025)
Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala
by: Rajapakse, Minuri, et al.
Published: (2026)
by: Rajapakse, Minuri, et al.
Published: (2026)
BERT-LID: Leveraging BERT to Improve Spoken Language Identification
by: Nie, Yuting, et al.
Published: (2022)
by: Nie, Yuting, et al.
Published: (2022)
On Entity Identification in Language Models
by: Sakata, Masaki, et al.
Published: (2025)
by: Sakata, Masaki, et al.
Published: (2025)
Script-Agnostic Language Identification
by: Agarwal, Milind, et al.
Published: (2024)
by: Agarwal, Milind, et al.
Published: (2024)
Geographically-Informed Language Identification
by: Dunn, Jonathan, et al.
Published: (2024)
by: Dunn, Jonathan, et al.
Published: (2024)
ILID: Native Script Language Identification for Indian Languages
by: Ingle, Yash, et al.
Published: (2025)
by: Ingle, Yash, et al.
Published: (2025)
GlotLID: Language Identification for Low-Resource Languages
by: Kargaran, Amir Hossein, et al.
Published: (2023)
by: Kargaran, Amir Hossein, et al.
Published: (2023)
Robust Language Identification for Romansh Varieties
by: Model, Charlotte, et al.
Published: (2026)
by: Model, Charlotte, et al.
Published: (2026)
The Hrunting of AI: Where and How to Improve English Dialectal Fairness
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
Blinded Multi-Rater Comparative Evaluation of a Large Language Model and Clinician-Authored Responses in CGM-Informed Diabetes Counseling
by: Guo, Zhijun, et al.
Published: (2026)
by: Guo, Zhijun, et al.
Published: (2026)
Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation
by: Gurgurov, Daniil, et al.
Published: (2025)
by: Gurgurov, Daniil, et al.
Published: (2025)
KréyoLID From Language Identification Towards Language Mining
by: Dent, Rasul, et al.
Published: (2025)
by: Dent, Rasul, et al.
Published: (2025)
Turkish Native Language Identification V2
by: Uluslu, Ahmet Yavuz, et al.
Published: (2023)
by: Uluslu, Ahmet Yavuz, et al.
Published: (2023)
Stylomech: Unveiling Authorship via Computational Stylometry in English and Romanized Sinhala
by: Faumi, Nabeelah, et al.
Published: (2025)
by: Faumi, Nabeelah, et al.
Published: (2025)
Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
by: Gharami, Kanchon, et al.
Published: (2025)
by: Gharami, Kanchon, et al.
Published: (2025)
Romanized to Native Malayalam Script Transliteration Using an Encoder-Decoder Framework
by: Baiju, Bajiyo, et al.
Published: (2024)
by: Baiju, Bajiyo, et al.
Published: (2024)
Leveraging Open-Source Large Language Models for Native Language Identification
by: Ng, Yee Man, et al.
Published: (2024)
by: Ng, Yee Man, et al.
Published: (2024)
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs
by: Mukherjee, Sagnik, et al.
Published: (2025)
by: Mukherjee, Sagnik, et al.
Published: (2025)
Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns
by: DuSell, Brian, et al.
Published: (2023)
by: DuSell, Brian, et al.
Published: (2023)
Similar Items
-
Graphemic Normalization of the Perso-Arabic Script
by: Doctor, Raiomond, et al.
Published: (2022) -
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023) -
Paradigm Completion for Derivational Morphology
by: Cotterell, Ryan, et al.
Published: (2017) -
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages
by: Ruder, Sebastian, et al.
Published: (2023) -
Exploring and Improving Drafts in Blockwise Parallel Decoding
by: Kim, Taehyeon, et al.
Published: (2024)