Graphemic Normalization of the Perso-Arabic Script
Fuente:
arXiv
Saved in:
| Main Authors: | Doctor, Raiomond, Gutkin, Alexander, Johny, Cibu, Roark, Brian, Sproat, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023)
by: Gutkin, Alexander, et al.
Published: (2023)
A Semantic Approach to Negation Detection and Word Disambiguation with Natural Language Processing
by: Okpala, Izunna, et al.
Published: (2023)
by: Okpala, Izunna, et al.
Published: (2023)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
by: Hossain, Ariyan, et al.
Published: (2025)
by: Hossain, Ariyan, et al.
Published: (2025)
HiPS: Hierarchical PDF Segmentation of Textbooks
by: Wehnert, Sabine, et al.
Published: (2025)
by: Wehnert, Sabine, et al.
Published: (2025)
The Need for Guardrails with Large Language Models in Medical Safety-Critical Settings: An Artificial Intelligence Application in the Pharmacovigilance Ecosystem
by: Hakim, Joe B, et al.
Published: (2024)
by: Hakim, Joe B, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
"The Data Says Otherwise"-Towards Automated Fact-checking and Communication of Data Claims
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
by: Bertina, Abbas, et al.
Published: (2025)
by: Bertina, Abbas, et al.
Published: (2025)
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
by: Verhoeff, Tom
Published: (2026)
by: Verhoeff, Tom
Published: (2026)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
by: Li, Lingxi, et al.
Published: (2024)
by: Li, Lingxi, et al.
Published: (2024)
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
by: Whittaker, Edward, et al.
Published: (2024)
by: Whittaker, Edward, et al.
Published: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing Process
by: Reza, Mohi, et al.
Published: (2025)
by: Reza, Mohi, et al.
Published: (2025)
Towards Full Authorship with AI: Supporting Revision with AI-Generated Views
by: Kim, Jiho, et al.
Published: (2024)
by: Kim, Jiho, et al.
Published: (2024)
Towards Conditioning Clinical Text Generation for User Control
by: Koraş, Osman Alperen, et al.
Published: (2025)
by: Koraş, Osman Alperen, et al.
Published: (2025)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
by: Han, Lifeng, et al.
Published: (2016)
by: Han, Lifeng, et al.
Published: (2016)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
by: Zhang, Luyan, et al.
Published: (2025)
by: Zhang, Luyan, et al.
Published: (2025)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
by: Zhu, Qian, et al.
Published: (2026)
by: Zhu, Qian, et al.
Published: (2026)
Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts
by: Glazkova, Anna, et al.
Published: (2024)
by: Glazkova, Anna, et al.
Published: (2024)
Normalization of Lithuanian Text Using Regular Expressions
by: Kasparaitis, Pijus
Published: (2023)
by: Kasparaitis, Pijus
Published: (2023)
S2Doc -- Spatial-Semantic Document Format
by: Kempf, Sebastian, et al.
Published: (2025)
by: Kempf, Sebastian, et al.
Published: (2025)
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025)
by: Doghmash, Salam Thabet, et al.
Published: (2025)
Robustness of Large Language Models to Perturbations in Text
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Dialect Normalization using Large Language Models and Morphological Rules
by: Dimakis, Antonios, et al.
Published: (2025)
by: Dimakis, Antonios, et al.
Published: (2025)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
by: Banthia, Saumya, et al.
Published: (2020)
by: Banthia, Saumya, et al.
Published: (2020)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
by: Zmanovskii, Nikita
Published: (2025)
by: Zmanovskii, Nikita
Published: (2025)
Breaking the HISCO Barrier: Automatic Occupational Standardization with OccCANINE
by: Dahl, Christian Møller, et al.
Published: (2024)
by: Dahl, Christian Møller, et al.
Published: (2024)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
by: Tu, Songjun, et al.
Published: (2026)
by: Tu, Songjun, et al.
Published: (2026)
Low-Resource Court Judgment Summarization for Common Law Systems
by: Liu, Shuaiqi, et al.
Published: (2024)
by: Liu, Shuaiqi, et al.
Published: (2024)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024)
by: Hashemi, Helia, et al.
Published: (2024)
Low-resource neural machine translation with morphological modeling
by: Nzeyimana, Antoine
Published: (2024)
by: Nzeyimana, Antoine
Published: (2024)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
by: Nzeyimana, Antoine, et al.
Published: (2025)
by: Nzeyimana, Antoine, et al.
Published: (2025)
MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
by: Wu, Yuexin, et al.
Published: (2025)
by: Wu, Yuexin, et al.
Published: (2025)
DROID: Dual Representation for Out-of-Scope Intent Detection
by: Rashwan, Wael, et al.
Published: (2025)
by: Rashwan, Wael, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Similar Items
-
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023) -
A Semantic Approach to Negation Detection and Word Disambiguation with Natural Language Processing
by: Okpala, Izunna, et al.
Published: (2023) -
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
by: Hossain, Ariyan, et al.
Published: (2025) -
HiPS: Hierarchical PDF Segmentation of Textbooks
by: Wehnert, Sabine, et al.
Published: (2025) -
The Need for Guardrails with Large Language Models in Medical Safety-Critical Settings: An Artificial Intelligence Application in the Pharmacovigilance Ecosystem
by: Hakim, Joe B, et al.
Published: (2024)