Faster Superword Tokenization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schmidt, Craig W., Tanner, Chris, Pinter, Yuval |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
von: Reddy, Varshini, et al.
Veröffentlicht: (2025)
von: Reddy, Varshini, et al.
Veröffentlicht: (2025)
Tokenization with Split Trees
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
The Effect of Scripts and Formats on LLM Numeracy
von: Reddy, Varshini, et al.
Veröffentlicht: (2026)
von: Reddy, Varshini, et al.
Veröffentlicht: (2026)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
Which Pieces Does Unigram Tokenization Really Need?
von: Land, Sander, et al.
Veröffentlicht: (2025)
von: Land, Sander, et al.
Veröffentlicht: (2025)
Splintering Nonconcatenative Languages for Better Tokenization
von: Gazit, Bar, et al.
Veröffentlicht: (2025)
von: Gazit, Bar, et al.
Veröffentlicht: (2025)
Protecting Privacy in Classifiers by Token Manipulation
von: Harel, Re'em, et al.
Veröffentlicht: (2024)
von: Harel, Re'em, et al.
Veröffentlicht: (2024)
Token-Level Privacy in Large Language Models
von: Harel, Re'em, et al.
Veröffentlicht: (2025)
von: Harel, Re'em, et al.
Veröffentlicht: (2025)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
The Degree of Language Diacriticity and Its Effect on Tasks
von: Cohen, Adi, et al.
Veröffentlicht: (2026)
von: Cohen, Adi, et al.
Veröffentlicht: (2026)
Probing Subphonemes in Morphology Models
von: Astrach, Gal, et al.
Veröffentlicht: (2025)
von: Astrach, Gal, et al.
Veröffentlicht: (2025)
Hebrew Diacritics Restoration using Visual Representation
von: Elboher, Yair, et al.
Veröffentlicht: (2025)
von: Elboher, Yair, et al.
Veröffentlicht: (2025)
Don't Touch My Diacritics
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
von: Cherf, Carinne, et al.
Veröffentlicht: (2024)
von: Cherf, Carinne, et al.
Veröffentlicht: (2024)
Information Types in Product Reviews
von: Shapira, Ori, et al.
Veröffentlicht: (2025)
von: Shapira, Ori, et al.
Veröffentlicht: (2025)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
von: Lai, Viet Dac, et al.
Veröffentlicht: (2024)
von: Lai, Viet Dac, et al.
Veröffentlicht: (2024)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token
von: Zychlinski, Shaked, et al.
Veröffentlicht: (2025)
von: Zychlinski, Shaked, et al.
Veröffentlicht: (2025)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
From Tokens to Words: On the Inner Lexicon of LLMs
von: Kaplan, Guy, et al.
Veröffentlicht: (2024)
von: Kaplan, Guy, et al.
Veröffentlicht: (2024)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
von: Han, Chao, et al.
Veröffentlicht: (2025)
von: Han, Chao, et al.
Veröffentlicht: (2025)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
von: Kaplan, Guy, et al.
Veröffentlicht: (2025)
von: Kaplan, Guy, et al.
Veröffentlicht: (2025)
OMPar: Automatic Parallelization with AI-Driven Source-to-Source Compilation
von: Kadosh, Tal, et al.
Veröffentlicht: (2024)
von: Kadosh, Tal, et al.
Veröffentlicht: (2024)
Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization
von: Balde, Gunjan, et al.
Veröffentlicht: (2026)
von: Balde, Gunjan, et al.
Veröffentlicht: (2026)
Morphologically-Informed Tokenizers for Languages with Non-Concatenative Morphology: A case study of Yoloxóchtil Mixtec ASR
von: Crawford, Chris
Veröffentlicht: (2025)
von: Crawford, Chris
Veröffentlicht: (2025)
Leveraging NTPs for Efficient Hallucination Detection in VLMs
von: Azachi, Ofir, et al.
Veröffentlicht: (2025)
von: Azachi, Ofir, et al.
Veröffentlicht: (2025)
Cost-Efficient Estimation of General Abilities Across Benchmarks
von: Krumdick, Michael, et al.
Veröffentlicht: (2026)
von: Krumdick, Michael, et al.
Veröffentlicht: (2026)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
von: Krumdick, Michael, et al.
Veröffentlicht: (2025)
von: Krumdick, Michael, et al.
Veröffentlicht: (2025)
Byte BPE Tokenization as an Inverse string Homomorphism
von: Geng, Saibo, et al.
Veröffentlicht: (2024)
von: Geng, Saibo, et al.
Veröffentlicht: (2024)
Incorporating Token Usage into Prompting Strategy Evaluation
von: Sypherd, Chris, et al.
Veröffentlicht: (2025)
von: Sypherd, Chris, et al.
Veröffentlicht: (2025)
Hypernym Mercury: Token Optimization Through Semantic Field Constriction And Reconstruction From Hypernyms. A New Text Compression Method
von: Forrester, Chris, et al.
Veröffentlicht: (2025)
von: Forrester, Chris, et al.
Veröffentlicht: (2025)
Project MOSLA: Recording Every Moment of Second Language Acquisition
von: Hagiwara, Masato, et al.
Veröffentlicht: (2024)
von: Hagiwara, Masato, et al.
Veröffentlicht: (2024)
Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
von: Koncel-Kedziorski, Rik, et al.
Veröffentlicht: (2023)
von: Koncel-Kedziorski, Rik, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
von: Uzan, Omri, et al.
Veröffentlicht: (2024) -
How Much is Enough? The Diminishing Returns of Tokenization Training Data
von: Reddy, Varshini, et al.
Veröffentlicht: (2025) -
Tokenization with Split Trees
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026) -
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025) -
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)