Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study
Fuente:
arXiv
Saved in:
| Main Authors: | Nadas, Mihai, Diosan, Laura |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic Data Generation Using Large Language Models: Advances in Text and Code
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction
by: Nadas, Mihai Dan, et al.
Published: (2026)
by: Nadas, Mihai Dan, et al.
Published: (2026)
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
by: Masala, Mihai, et al.
Published: (2026)
by: Masala, Mihai, et al.
Published: (2026)
Hebrew Diacritics Restoration using Visual Representation
by: Elboher, Yair, et al.
Published: (2025)
by: Elboher, Yair, et al.
Published: (2025)
Automatic Restoration of Diacritics for Speech Data Sets
by: Shatnawi, Sara, et al.
Published: (2023)
by: Shatnawi, Sara, et al.
Published: (2023)
Diacritic Restoration for Low-Resource Indigenous Languages: Case Study with Bribri and Cook Islands Māori
by: Coto-Solano, Rolando, et al.
Published: (2025)
by: Coto-Solano, Rolando, et al.
Published: (2025)
Corpus-Based Approaches to Igbo Diacritic Restoration
by: Ezeani, Ignatius
Published: (2026)
by: Ezeani, Ignatius
Published: (2026)
LLMic: Romanian Foundation Language Model
by: Bădoiu, Vlad-Andrei, et al.
Published: (2025)
by: Bădoiu, Vlad-Andrei, et al.
Published: (2025)
Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations
by: Ghannam, Ahmad, et al.
Published: (2025)
by: Ghannam, Ahmad, et al.
Published: (2025)
Interplay of Machine Translation, Diacritics, and Diacritization
by: Chen, Wei-Rui, et al.
Published: (2024)
by: Chen, Wei-Rui, et al.
Published: (2024)
Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization
by: Elgamal, Salman, et al.
Published: (2024)
by: Elgamal, Salman, et al.
Published: (2024)
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
Improving Legal Judgement Prediction in Romanian with Long Text Encoders
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
Neural Grammatical Error Correction for Romanian
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text
by: Olawole, Akindele Michael, et al.
Published: (2024)
by: Olawole, Akindele Michael, et al.
Published: (2024)
The Degree of Language Diacriticity and Its Effect on Tasks
by: Cohen, Adi, et al.
Published: (2026)
by: Cohen, Adi, et al.
Published: (2026)
FuLG: 150B Romanian Corpus for Language Model Pretraining
by: Bădoiu, Vlad-Andrei, et al.
Published: (2024)
by: Bădoiu, Vlad-Andrei, et al.
Published: (2024)
Sadeed: Advancing Arabic Diacritization Through Small Language Model
by: Aldallal, Zeina, et al.
Published: (2025)
by: Aldallal, Zeina, et al.
Published: (2025)
Don't Touch My Diacritics
by: Gorman, Kyle, et al.
Published: (2024)
by: Gorman, Kyle, et al.
Published: (2024)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
by: Roth, Amit, et al.
Published: (2024)
by: Roth, Amit, et al.
Published: (2024)
Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need
by: Skiredj, Abderrahman, et al.
Published: (2024)
by: Skiredj, Abderrahman, et al.
Published: (2024)
D-Nikud: Enhancing Hebrew Diacritization with LSTM and Pretrained Models
by: Rosenthal, Adi, et al.
Published: (2024)
by: Rosenthal, Adi, et al.
Published: (2024)
A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
by: Rogoz, Ana-Cristina, et al.
Published: (2025)
by: Rogoz, Ana-Cristina, et al.
Published: (2025)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
by: Negoita, Vlad, et al.
Published: (2025)
by: Negoita, Vlad, et al.
Published: (2025)
Fine-Tuning Large Language Models for Scientific Text Classification: A Comparative Study
by: Rostam, Zhyar Rzgar K, et al.
Published: (2024)
by: Rostam, Zhyar Rzgar K, et al.
Published: (2024)
Proper Noun Diacritization for Arabic Wikipedia: A Benchmark Dataset
by: Bondok, Rawan, et al.
Published: (2025)
by: Bondok, Rawan, et al.
Published: (2025)
More Data, Fewer Diacritics: Scaling Arabic TTS
by: Musleh, Ahmed, et al.
Published: (2026)
by: Musleh, Ahmed, et al.
Published: (2026)
OpenLLM-Ro -- Technical Report on Open-source Romanian LLMs
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
by: Do, Thao, et al.
Published: (2024)
by: Do, Thao, et al.
Published: (2024)
AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
by: Afzal, Anum, et al.
Published: (2024)
by: Afzal, Anum, et al.
Published: (2024)
A Context-Contrastive Inference Approach To Partial Diacritization
by: ElNokrashy, Muhammad, et al.
Published: (2024)
by: ElNokrashy, Muhammad, et al.
Published: (2024)
RoQLlama: A Lightweight Romanian Adapted Language Model
by: Dima, George-Andrei, et al.
Published: (2024)
by: Dima, George-Andrei, et al.
Published: (2024)
Irony Detection in Urdu Text: A Comparative Study Using Machine Learning Models and Large Language Models
by: Ahmad, Fiaz, et al.
Published: (2025)
by: Ahmad, Fiaz, et al.
Published: (2025)
A Comparative Evaluation of Large Language Models for Persian Sentiment Analysis and Emotion Detection in Social Media Texts
by: Tohidi, Kian, et al.
Published: (2025)
by: Tohidi, Kian, et al.
Published: (2025)
PsihoRo: Depression and Anxiety Romanian Text Corpus
by: Ciobotaru, Alexandra, et al.
Published: (2026)
by: Ciobotaru, Alexandra, et al.
Published: (2026)
"Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
A Cross-Lingual Analysis of Bias in Large Language Models Using Romanian History
by: Cocu, Matei-Iulian, et al.
Published: (2025)
by: Cocu, Matei-Iulian, et al.
Published: (2025)
Exploring Large Language Models for Translating Romanian Computational Problems into English
by: Dumitran, Adrian Marius, et al.
Published: (2025)
by: Dumitran, Adrian Marius, et al.
Published: (2025)
Similar Items
-
Synthetic Data Generation Using Large Language Models: Advances in Text and Code
by: Nadas, Mihai, et al.
Published: (2025) -
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
by: Nadas, Mihai, et al.
Published: (2025) -
TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction
by: Nadas, Mihai Dan, et al.
Published: (2026) -
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
by: Nadas, Mihai, et al.
Published: (2025) -
"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
by: Masala, Mihai, et al.
Published: (2026)