An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Cognetta, Marco, Hiraoka, Tatsuya, Okazaki, Naoaki, Sennrich, Rico, Pinter, Yuval |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Tokenization as Finite-State Transduction
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
by: Cherf, Carinne, et al.
Published: (2024)
by: Cherf, Carinne, et al.
Published: (2024)
Two Counterexamples to Tokenization and the Noiseless Channel
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Decoding-Free Sampling Strategies for LLM Marginalization
by: Pohl, David, et al.
Published: (2025)
by: Pohl, David, et al.
Published: (2025)
Machine Translation Models are Zero-Shot Detectors of Translation Direction
by: Wastl, Michelle, et al.
Published: (2024)
by: Wastl, Michelle, et al.
Published: (2024)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
by: Wu, Guojun, et al.
Published: (2024)
by: Wu, Guojun, et al.
Published: (2024)
Mitigating Hallucinations and Off-target Machine Translation with Source-Contrastive and Language-Contrastive Decoding
by: Sennrich, Rico, et al.
Published: (2023)
by: Sennrich, Rico, et al.
Published: (2023)
Investigating Multi-Pivot Ensembling with Massively Multilingual Machine Translation Models
by: Mohammadshahi, Alireza, et al.
Published: (2023)
by: Mohammadshahi, Alireza, et al.
Published: (2023)
Source-primed Multi-turn Conversation Helps Large Language Models Translate Documents
by: Hu, Hanxu, et al.
Published: (2025)
by: Hu, Hanxu, et al.
Published: (2025)
Tokenization Preference for Human and Machine Learning Model: An Annotation Study
by: Hiraoka, Tatsuya, et al.
Published: (2023)
by: Hiraoka, Tatsuya, et al.
Published: (2023)
Tutorial: $φ$-Transductions in OpenFst via the Gallic Semiring
by: Cognetta, Marco, et al.
Published: (2025)
by: Cognetta, Marco, et al.
Published: (2025)
Linear-time Minimum Bayes Risk Decoding with Reference Aggregation
by: Vamvas, Jannis, et al.
Published: (2024)
by: Vamvas, Jannis, et al.
Published: (2024)
Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
by: Semenov, Kirill, et al.
Published: (2025)
by: Semenov, Kirill, et al.
Published: (2025)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
by: Moghe, Nikita, et al.
Published: (2024)
by: Moghe, Nikita, et al.
Published: (2024)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
by: Ichinose, Tatsuya, et al.
Published: (2026)
by: Ichinose, Tatsuya, et al.
Published: (2026)
Don't Touch My Diacritics
by: Gorman, Kyle, et al.
Published: (2024)
by: Gorman, Kyle, et al.
Published: (2024)
Probing Subphonemes in Morphology Models
by: Astrach, Gal, et al.
Published: (2025)
by: Astrach, Gal, et al.
Published: (2025)
Hebrew Diacritics Restoration using Visual Representation
by: Elboher, Yair, et al.
Published: (2025)
by: Elboher, Yair, et al.
Published: (2025)
Information Types in Product Reviews
by: Shapira, Ori, et al.
Published: (2025)
by: Shapira, Ori, et al.
Published: (2025)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
The Degree of Language Diacriticity and Its Effect on Tasks
by: Cohen, Adi, et al.
Published: (2026)
by: Cohen, Adi, et al.
Published: (2026)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024)
by: Chizhov, Pavel, et al.
Published: (2024)
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer
by: Ma, Youmi, et al.
Published: (2024)
by: Ma, Youmi, et al.
Published: (2024)
Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025)
by: Hiraoka, Tatsuya, et al.
Published: (2025)
Modular Adaptation of Multilingual Encoders to Written Swiss German Dialect
by: Vamvas, Jannis, et al.
Published: (2024)
by: Vamvas, Jannis, et al.
Published: (2024)
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
by: Michail, Andrianos, et al.
Published: (2025)
by: Michail, Andrianos, et al.
Published: (2025)
SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents
by: Wastl, Michelle, et al.
Published: (2025)
by: Wastl, Michelle, et al.
Published: (2025)
Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
by: Kew, Tannon, et al.
Published: (2023)
by: Kew, Tannon, et al.
Published: (2023)
SwissBERT: The Multilingual Language Model for Switzerland
by: Vamvas, Jannis, et al.
Published: (2023)
by: Vamvas, Jannis, et al.
Published: (2023)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
Similar Items
-
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025) -
Tokenization as Finite-State Transduction
by: Cognetta, Marco, et al.
Published: (2024) -
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024) -
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024) -
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
by: Cherf, Carinne, et al.
Published: (2024)