AdaptBPE: From General Purpose to Specialized Tokenizers
Fuente:
arXiv
Guardado en:
| Autores principales: | Liyanage, Vijini, Yvon, François |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BlockBPE: Parallel BPE Tokenization
por: You, Amos
Publicado: (2025)
por: You, Amos
Publicado: (2025)
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
por: Dolga, Rares, et al.
Publicado: (2025)
por: Dolga, Rares, et al.
Publicado: (2025)
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
por: Sun, Yike, et al.
Publicado: (2026)
por: Sun, Yike, et al.
Publicado: (2026)
Batching BPE Tokenization Merges
por: Morgan, Alexander P.
Publicado: (2024)
por: Morgan, Alexander P.
Publicado: (2024)
Constructing a BPE Tokenization DFA
por: Berglund, Martin, et al.
Publicado: (2024)
por: Berglund, Martin, et al.
Publicado: (2024)
Byte BPE Tokenization as an Inverse string Homomorphism
por: Geng, Saibo, et al.
Publicado: (2024)
por: Geng, Saibo, et al.
Publicado: (2024)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
por: Chizhov, Pavel, et al.
Publicado: (2024)
por: Chizhov, Pavel, et al.
Publicado: (2024)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
por: Patwary, Firoj Ahmmed, et al.
Publicado: (2025)
por: Patwary, Firoj Ahmmed, et al.
Publicado: (2025)
Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
por: Balde, Gunjan, et al.
Publicado: (2024)
por: Balde, Gunjan, et al.
Publicado: (2024)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
por: Kadamba, Venu Gopal, et al.
Publicado: (2026)
por: Kadamba, Venu Gopal, et al.
Publicado: (2026)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
por: Vemula, Saketh Reddy, et al.
Publicado: (2025)
por: Vemula, Saketh Reddy, et al.
Publicado: (2025)
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
por: Lerner, Paul, et al.
Publicado: (2025)
por: Lerner, Paul, et al.
Publicado: (2025)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
por: Hayase, Jonathan, et al.
Publicado: (2024)
por: Hayase, Jonathan, et al.
Publicado: (2024)
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
por: Lian, Haoran, et al.
Publicado: (2024)
por: Lian, Haoran, et al.
Publicado: (2024)
Bit-level BPE: Below the byte boundary
por: Moon, Sangwhan, et al.
Publicado: (2025)
por: Moon, Sangwhan, et al.
Publicado: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
por: Parra, Iñigo
Publicado: (2024)
por: Parra, Iñigo
Publicado: (2024)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
por: Cognetta, Marco, et al.
Publicado: (2024)
por: Cognetta, Marco, et al.
Publicado: (2024)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
por: Asgari, Ehsaneddin, et al.
Publicado: (2025)
por: Asgari, Ehsaneddin, et al.
Publicado: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
por: Labat, Léo, et al.
Publicado: (2026)
por: Labat, Léo, et al.
Publicado: (2026)
How Sampling Affects the Detectability of Machine-written texts: A Comprehensive Study
por: Dubois, Matthieu, et al.
Publicado: (2025)
por: Dubois, Matthieu, et al.
Publicado: (2025)
Optimizing example selection for retrieval-augmented machine translation with translation memories
por: Bouthors, Maxime, et al.
Publicado: (2024)
por: Bouthors, Maxime, et al.
Publicado: (2024)
Retrieving Examples from Memory for Retrieval Augmented Neural Machine Translation: A Systematic Comparison
por: Bouthors, Maxime, et al.
Publicado: (2024)
por: Bouthors, Maxime, et al.
Publicado: (2024)
Prompting LLMs: Length Control for Isometric Machine Translation
por: Javorský, Dávid, et al.
Publicado: (2025)
por: Javorský, Dávid, et al.
Publicado: (2025)
Investigating Length Issues in Document-level Machine Translation
por: Peng, Ziqian, et al.
Publicado: (2024)
por: Peng, Ziqian, et al.
Publicado: (2024)
MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
por: Javorský, Dávid, et al.
Publicado: (2025)
por: Javorský, Dávid, et al.
Publicado: (2025)
MOSAIC: Multiple Observers Spotting AI Content
por: Dubois, Matthieu, et al.
Publicado: (2024)
por: Dubois, Matthieu, et al.
Publicado: (2024)
Subasa - Adapting Language Models for Low-resourced Offensive Language Detection in Sinhala
por: Haturusinghe, Shanilka, et al.
Publicado: (2025)
por: Haturusinghe, Shanilka, et al.
Publicado: (2025)
SuperBPE: Space Travel for Language Models
por: Liu, Alisa, et al.
Publicado: (2025)
por: Liu, Alisa, et al.
Publicado: (2025)
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
por: Land, Sander, et al.
Publicado: (2025)
por: Land, Sander, et al.
Publicado: (2025)
Adapting General-Purpose Embedding Models to Private Datasets Using Keyword-based Retrieval
por: Wei, Yubai, et al.
Publicado: (2025)
por: Wei, Yubai, et al.
Publicado: (2025)
Tag-LLM: Repurposing General-Purpose LLMs for Specialized Domains
por: Shen, Junhong, et al.
Publicado: (2024)
por: Shen, Junhong, et al.
Publicado: (2024)
MaskLID: Code-Switching Language Identification through Iterative Masking
por: Kargaran, Amir Hossein, et al.
Publicado: (2024)
por: Kargaran, Amir Hossein, et al.
Publicado: (2024)
GlotScript: A Resource and Tool for Low Resource Writing System Identification
por: Kargaran, Amir Hossein, et al.
Publicado: (2023)
por: Kargaran, Amir Hossein, et al.
Publicado: (2023)
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models
por: Sawada, Tomohiro, et al.
Publicado: (2025)
por: Sawada, Tomohiro, et al.
Publicado: (2025)
On the Entity-Level Alignment in Crosslingual Consistency
por: Liu, Yihong, et al.
Publicado: (2025)
por: Liu, Yihong, et al.
Publicado: (2025)
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
por: Bouthors, Maxime, et al.
Publicado: (2025)
por: Bouthors, Maxime, et al.
Publicado: (2025)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
por: Kargaran, Amir Hossein, et al.
Publicado: (2024)
por: Kargaran, Amir Hossein, et al.
Publicado: (2024)
GlotLID: Language Identification for Low-Resource Languages
por: Kargaran, Amir Hossein, et al.
Publicado: (2023)
por: Kargaran, Amir Hossein, et al.
Publicado: (2023)
How Programming Concepts and Neurons Are Shared in Code Language Models
por: Kargaran, Amir Hossein, et al.
Publicado: (2025)
por: Kargaran, Amir Hossein, et al.
Publicado: (2025)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
Ejemplares similares
-
BlockBPE: Parallel BPE Tokenization
por: You, Amos
Publicado: (2025) -
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
por: Dolga, Rares, et al.
Publicado: (2025) -
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
por: Sun, Yike, et al.
Publicado: (2026) -
Batching BPE Tokenization Merges
por: Morgan, Alexander P.
Publicado: (2024) -
Constructing a BPE Tokenization DFA
por: Berglund, Martin, et al.
Publicado: (2024)