Don't Touch My Diacritics
Fuente:
arXiv
Saved in:
| Main Authors: | Gorman, Kyle, Pinter, Yuval |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hebrew Diacritics Restoration using Visual Representation
by: Elboher, Yair, et al.
Published: (2025)
by: Elboher, Yair, et al.
Published: (2025)
The Degree of Language Diacriticity and Its Effect on Tasks
by: Cohen, Adi, et al.
Published: (2026)
by: Cohen, Adi, et al.
Published: (2026)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
by: Cherf, Carinne, et al.
Published: (2024)
by: Cherf, Carinne, et al.
Published: (2024)
Probing Subphonemes in Morphology Models
by: Astrach, Gal, et al.
Published: (2025)
by: Astrach, Gal, et al.
Published: (2025)
Information Types in Product Reviews
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
Interplay of Machine Translation, Diacritics, and Diacritization
by: Chen, Wei-Rui, et al.
Published: (2024)
by: Chen, Wei-Rui, et al.
Published: (2024)
Arabic Diacritics in the Wild: Exploiting Opportunities for Improved Diacritization
by: Elgamal, Salman, et al.
Published: (2024)
by: Elgamal, Salman, et al.
Published: (2024)
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
Don't Change My View: Ideological Bias Auditing in Large Language Models
by: Kröger, Paul, et al.
Published: (2025)
by: Kröger, Paul, et al.
Published: (2025)
Don't Pay Attention
by: Hammoud, Mohammad, et al.
Published: (2025)
by: Hammoud, Mohammad, et al.
Published: (2025)
Faster Superword Tokenization
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
A* shortest string decoding for non-idempotent semirings
by: Gorman, Kyle, et al.
Published: (2022)
by: Gorman, Kyle, et al.
Published: (2022)
Splintering Nonconcatenative Languages for Better Tokenization
by: Gazit, Bar, et al.
Published: (2025)
by: Gazit, Bar, et al.
Published: (2025)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
Protecting Privacy in Classifiers by Token Manipulation
by: Harel, Re'em, et al.
Published: (2024)
by: Harel, Re'em, et al.
Published: (2024)
Token-Level Privacy in Large Language Models
by: Harel, Re'em, et al.
Published: (2025)
by: Harel, Re'em, et al.
Published: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024)
by: Zhou, Yukai, et al.
Published: (2024)
Don't Throw Away Your Pretrained Model
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
Automatic Restoration of Diacritics for Speech Data Sets
by: Shatnawi, Sara, et al.
Published: (2023)
by: Shatnawi, Sara, et al.
Published: (2023)
More Data, Fewer Diacritics: Scaling Arabic TTS
by: Musleh, Ahmed, et al.
Published: (2026)
by: Musleh, Ahmed, et al.
Published: (2026)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
by: Uzan, Omri, et al.
Published: (2024)
by: Uzan, Omri, et al.
Published: (2024)
Corpus-Based Approaches to Igbo Diacritic Restoration
by: Ezeani, Ignatius
Published: (2026)
by: Ezeani, Ignatius
Published: (2026)
Hatevolution: What Static Benchmarks Don't Tell Us
by: Di Bonaventura, Chiara, et al.
Published: (2025)
by: Di Bonaventura, Chiara, et al.
Published: (2025)
Think, But Don't Overthink: Reproducing Recursive Language Models
by: Wang, Daren
Published: (2026)
by: Wang, Daren
Published: (2026)
YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text
by: Olawole, Akindele Michael, et al.
Published: (2024)
by: Olawole, Akindele Michael, et al.
Published: (2024)
D-Nikud: Enhancing Hebrew Diacritization with LSTM and Pretrained Models
by: Rosenthal, Adi, et al.
Published: (2024)
by: Rosenthal, Adi, et al.
Published: (2024)
Proper Noun Diacritization for Arabic Wikipedia: A Benchmark Dataset
by: Bondok, Rawan, et al.
Published: (2025)
by: Bondok, Rawan, et al.
Published: (2025)
Reasoning Models Reason Well, Until They Don't
by: Rameshkumar, Revanth, et al.
Published: (2025)
by: Rameshkumar, Revanth, et al.
Published: (2025)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Don't Throw Away Data: Better Sequence Knowledge Distillation
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Don't Command, Cultivate: An Exploratory Study of System-2 Alignment
by: Wang, Yuhang, et al.
Published: (2024)
by: Wang, Yuhang, et al.
Published: (2024)
A Context-Contrastive Inference Approach To Partial Diacritization
by: ElNokrashy, Muhammad, et al.
Published: (2024)
by: ElNokrashy, Muhammad, et al.
Published: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
by: Reddy, Varshini, et al.
Published: (2025)
by: Reddy, Varshini, et al.
Published: (2025)
Language Models Don't Learn the Physical Manifestation of Language
by: Lee, Bruce W., et al.
Published: (2024)
by: Lee, Bruce W., et al.
Published: (2024)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Don't Walk the Line: Boundary Guidance for Filtered Generation
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That!
by: Bafna, Niyati, et al.
Published: (2024)
by: Bafna, Niyati, et al.
Published: (2024)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
Similar Items
-
Hebrew Diacritics Restoration using Visual Representation
by: Elboher, Yair, et al.
Published: (2025) -
The Degree of Language Diacriticity and Its Effect on Tasks
by: Cohen, Adi, et al.
Published: (2026) -
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
by: Cherf, Carinne, et al.
Published: (2024) -
Probing Subphonemes in Morphology Models
by: Astrach, Gal, et al.
Published: (2025) -
Information Types in Product Reviews
by: Shapira, Ori, et al.
Published: (2025)