Greed is All You Need: An Evaluation of Tokenizer Inference Methods
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Uzan, Omri, Schmidt, Craig W., Tanner, Chris, Pinter, Yuval |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
Faster Superword Tokenization
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
von: Reddy, Varshini, et al.
Veröffentlicht: (2025)
von: Reddy, Varshini, et al.
Veröffentlicht: (2025)
Tokenization with Split Trees
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
Which Pieces Does Unigram Tokenization Really Need?
von: Land, Sander, et al.
Veröffentlicht: (2025)
von: Land, Sander, et al.
Veröffentlicht: (2025)
The Effect of Scripts and Formats on LLM Numeracy
von: Reddy, Varshini, et al.
Veröffentlicht: (2026)
von: Reddy, Varshini, et al.
Veröffentlicht: (2026)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
von: Cherf, Carinne, et al.
Veröffentlicht: (2024)
von: Cherf, Carinne, et al.
Veröffentlicht: (2024)
Not All Tokens Are What You Need In Thinking
von: Yuan, Hang, et al.
Veröffentlicht: (2025)
von: Yuan, Hang, et al.
Veröffentlicht: (2025)
Splintering Nonconcatenative Languages for Better Tokenization
von: Gazit, Bar, et al.
Veröffentlicht: (2025)
von: Gazit, Bar, et al.
Veröffentlicht: (2025)
Protecting Privacy in Classifiers by Token Manipulation
von: Harel, Re'em, et al.
Veröffentlicht: (2024)
von: Harel, Re'em, et al.
Veröffentlicht: (2024)
Token-Level Privacy in Large Language Models
von: Harel, Re'em, et al.
Veröffentlicht: (2025)
von: Harel, Re'em, et al.
Veröffentlicht: (2025)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
Rho-1: Not All Tokens Are What You Need
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
Don't Touch My Diacritics
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
Probing Subphonemes in Morphology Models
von: Astrach, Gal, et al.
Veröffentlicht: (2025)
von: Astrach, Gal, et al.
Veröffentlicht: (2025)
Hebrew Diacritics Restoration using Visual Representation
von: Elboher, Yair, et al.
Veröffentlicht: (2025)
von: Elboher, Yair, et al.
Veröffentlicht: (2025)
Information Types in Product Reviews
von: Shapira, Ori, et al.
Veröffentlicht: (2025)
von: Shapira, Ori, et al.
Veröffentlicht: (2025)
The Degree of Language Diacriticity and Its Effect on Tasks
von: Cohen, Adi, et al.
Veröffentlicht: (2026)
von: Cohen, Adi, et al.
Veröffentlicht: (2026)
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
von: Ji, Ke, et al.
Veröffentlicht: (2025)
von: Ji, Ke, et al.
Veröffentlicht: (2025)
Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need
von: Skiredj, Abderrahman, et al.
Veröffentlicht: (2024)
von: Skiredj, Abderrahman, et al.
Veröffentlicht: (2024)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
von: Lai, Viet Dac, et al.
Veröffentlicht: (2024)
von: Lai, Viet Dac, et al.
Veröffentlicht: (2024)
Document Optimization for Black-Box Retrieval via Reinforcement Learning
von: Uzan, Omri, et al.
Veröffentlicht: (2026)
von: Uzan, Omri, et al.
Veröffentlicht: (2026)
Contrast Is All You Need
von: Kilic, Burak, et al.
Veröffentlicht: (2023)
von: Kilic, Burak, et al.
Veröffentlicht: (2023)
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
Increasing the Thinking Budget is Not All You Need
von: Iacobacci, Ignacio, et al.
Veröffentlicht: (2025)
von: Iacobacci, Ignacio, et al.
Veröffentlicht: (2025)
Training on the Benchmark Is Not All You Need
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
Inference is All You Need: Self Example Retriever for Cross-domain Dialogue State Tracking with ChatGPT
von: Lee, Jihyun, et al.
Veröffentlicht: (2024)
von: Lee, Jihyun, et al.
Veröffentlicht: (2024)
Agents Are All You Need for LLM Unlearning
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need
von: Wang, Yang, et al.
Veröffentlicht: (2024)
von: Wang, Yang, et al.
Veröffentlicht: (2024)
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
von: Guo, Zhiyu, et al.
Veröffentlicht: (2024)
von: Guo, Zhiyu, et al.
Veröffentlicht: (2024)
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024)
von: Li, Junyou, et al.
Veröffentlicht: (2024)
Grimoire is All You Need for Enhancing Large Language Models
von: Chen, Ding, et al.
Veröffentlicht: (2024)
von: Chen, Ding, et al.
Veröffentlicht: (2024)
Addition is All You Need for Energy-efficient Language Models
von: Luo, Hongyin, et al.
Veröffentlicht: (2024)
von: Luo, Hongyin, et al.
Veröffentlicht: (2024)
Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
von: Lall, Supriya, et al.
Veröffentlicht: (2025)
von: Lall, Supriya, et al.
Veröffentlicht: (2025)
Lexicalization Is All You Need: Examining the Impact of Lexical Knowledge in a Compositional QALD System
von: Schmidt, David Maria, et al.
Veröffentlicht: (2024)
von: Schmidt, David Maria, et al.
Veröffentlicht: (2024)
Spontaneous Giving and Calculated Greed in Language Models
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
von: Uzan, Omri, et al.
Veröffentlicht: (2025) -
Faster Superword Tokenization
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026) -
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024) -
How Much is Enough? The Diminishing Returns of Tokenization Training Data
von: Reddy, Varshini, et al.
Veröffentlicht: (2025) -
Tokenization with Split Trees
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)