Tokenization with Split Trees
Fuente:
arXiv
Salvato in:
| Autori principali: | Schmidt, Craig W., Krumdick, Michael, Wiemerslage, Adam, Ebner, Seth, Reddy, Varshini, Pinter, Yuval, Tanner, Chris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Effect of Scripts and Formats on LLM Numeracy
di: Reddy, Varshini, et al.
Pubblicazione: (2026)
di: Reddy, Varshini, et al.
Pubblicazione: (2026)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
di: Reddy, Varshini, et al.
Pubblicazione: (2025)
di: Reddy, Varshini, et al.
Pubblicazione: (2025)
Cost-Efficient Estimation of General Abilities Across Benchmarks
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
di: Krumdick, Michael, et al.
Pubblicazione: (2025)
di: Krumdick, Michael, et al.
Pubblicazione: (2025)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
di: Schmidt, Craig W., et al.
Pubblicazione: (2025)
di: Schmidt, Craig W., et al.
Pubblicazione: (2025)
Faster Superword Tokenization
di: Schmidt, Craig W., et al.
Pubblicazione: (2026)
di: Schmidt, Craig W., et al.
Pubblicazione: (2026)
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
di: Uzan, Omri, et al.
Pubblicazione: (2024)
di: Uzan, Omri, et al.
Pubblicazione: (2024)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
di: Lai, Viet Dac, et al.
Pubblicazione: (2024)
di: Lai, Viet Dac, et al.
Pubblicazione: (2024)
On Finding Inconsistencies in Documents
di: Lovering, Charles J., et al.
Pubblicazione: (2025)
di: Lovering, Charles J., et al.
Pubblicazione: (2025)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
di: Koncel-Kedziorski, Rik, et al.
Pubblicazione: (2023)
di: Koncel-Kedziorski, Rik, et al.
Pubblicazione: (2023)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
DocFinQA: A Long-Context Financial Reasoning Dataset
di: Reddy, Varshini, et al.
Pubblicazione: (2024)
di: Reddy, Varshini, et al.
Pubblicazione: (2024)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
di: Uzan, Omri, et al.
Pubblicazione: (2025)
di: Uzan, Omri, et al.
Pubblicazione: (2025)
An Analysis of Multilingual FActScore
di: Vu, Kim Trong, et al.
Pubblicazione: (2024)
di: Vu, Kim Trong, et al.
Pubblicazione: (2024)
Which Pieces Does Unigram Tokenization Really Need?
di: Land, Sander, et al.
Pubblicazione: (2025)
di: Land, Sander, et al.
Pubblicazione: (2025)
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
di: Krumdick, Michael, et al.
Pubblicazione: (2026)
From Algebraic Word Problem to Program: A Formalized Approach
di: Wiemerslage, Adam, et al.
Pubblicazione: (2020)
di: Wiemerslage, Adam, et al.
Pubblicazione: (2020)
Splintering Nonconcatenative Languages for Better Tokenization
di: Gazit, Bar, et al.
Pubblicazione: (2025)
di: Gazit, Bar, et al.
Pubblicazione: (2025)
Protecting Privacy in Classifiers by Token Manipulation
di: Harel, Re'em, et al.
Pubblicazione: (2024)
di: Harel, Re'em, et al.
Pubblicazione: (2024)
Token-Level Privacy in Large Language Models
di: Harel, Re'em, et al.
Pubblicazione: (2025)
di: Harel, Re'em, et al.
Pubblicazione: (2025)
Language Model Probabilities are Not Calibrated in Numeric Contexts
di: Lovering, Charles, et al.
Pubblicazione: (2024)
di: Lovering, Charles, et al.
Pubblicazione: (2024)
The Degree of Language Diacriticity and Its Effect on Tasks
di: Cohen, Adi, et al.
Pubblicazione: (2026)
di: Cohen, Adi, et al.
Pubblicazione: (2026)
Probing Subphonemes in Morphology Models
di: Astrach, Gal, et al.
Pubblicazione: (2025)
di: Astrach, Gal, et al.
Pubblicazione: (2025)
Hebrew Diacritics Restoration using Visual Representation
di: Elboher, Yair, et al.
Pubblicazione: (2025)
di: Elboher, Yair, et al.
Pubblicazione: (2025)
Don't Touch My Diacritics
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
di: Cherf, Carinne, et al.
Pubblicazione: (2024)
di: Cherf, Carinne, et al.
Pubblicazione: (2024)
Information Types in Product Reviews
di: Shapira, Ori, et al.
Pubblicazione: (2025)
di: Shapira, Ori, et al.
Pubblicazione: (2025)
Improving Low-Resource Morphological Inflection via Self-Supervised Objectives
di: Wiemerslage, Adam, et al.
Pubblicazione: (2025)
di: Wiemerslage, Adam, et al.
Pubblicazione: (2025)
Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer
di: Ebrahimi, Abteen, et al.
Pubblicazione: (2025)
di: Ebrahimi, Abteen, et al.
Pubblicazione: (2025)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
di: Chang, Yapei, et al.
Pubblicazione: (2025)
di: Chang, Yapei, et al.
Pubblicazione: (2025)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
A Closer Look at Claim Decomposition
di: Wanner, Miriam, et al.
Pubblicazione: (2024)
di: Wanner, Miriam, et al.
Pubblicazione: (2024)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
di: Kaplan, Guy, et al.
Pubblicazione: (2025)
di: Kaplan, Guy, et al.
Pubblicazione: (2025)
Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token
di: Zychlinski, Shaked, et al.
Pubblicazione: (2025)
di: Zychlinski, Shaked, et al.
Pubblicazione: (2025)
From Tokens to Words: On the Inner Lexicon of LLMs
di: Kaplan, Guy, et al.
Pubblicazione: (2024)
di: Kaplan, Guy, et al.
Pubblicazione: (2024)
OMPar: Automatic Parallelization with AI-Driven Source-to-Source Compilation
di: Kadosh, Tal, et al.
Pubblicazione: (2024)
di: Kadosh, Tal, et al.
Pubblicazione: (2024)
Morphologically-Informed Tokenizers for Languages with Non-Concatenative Morphology: A case study of Yoloxóchtil Mixtec ASR
di: Crawford, Chris
Pubblicazione: (2025)
di: Crawford, Chris
Pubblicazione: (2025)
Documenti analoghi
-
The Effect of Scripts and Formats on LLM Numeracy
di: Reddy, Varshini, et al.
Pubblicazione: (2026) -
How Much is Enough? The Diminishing Returns of Tokenization Training Data
di: Reddy, Varshini, et al.
Pubblicazione: (2025) -
Cost-Efficient Estimation of General Abilities Across Benchmarks
di: Krumdick, Michael, et al.
Pubblicazione: (2026) -
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
di: Krumdick, Michael, et al.
Pubblicazione: (2025) -
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
di: Schmidt, Craig W., et al.
Pubblicazione: (2025)