Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
Fuente:
arXiv
Salvato in:
| Autori principali: | Goldman, Omer, Caciularu, Avi, Eyal, Matan, Cao, Kris, Szpektor, Idan, Tsarfaty, Reut |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
di: Shaham, Uri, et al.
Pubblicazione: (2024)
di: Shaham, Uri, et al.
Pubblicazione: (2024)
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
di: Mor-Lan, Guy, et al.
Pubblicazione: (2026)
di: Mor-Lan, Guy, et al.
Pubblicazione: (2026)
ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer
di: Goldman, Omer, et al.
Pubblicazione: (2025)
di: Goldman, Omer, et al.
Pubblicazione: (2025)
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
di: Greenfeld, Refael Shaked, et al.
Pubblicazione: (2026)
di: Greenfeld, Refael Shaked, et al.
Pubblicazione: (2026)
A Truly Joint Neural Architecture for Segmentation and Parsing
di: Levi, Danit Yshaayahu, et al.
Pubblicazione: (2024)
di: Levi, Danit Yshaayahu, et al.
Pubblicazione: (2024)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
di: Cattan, Arie, et al.
Pubblicazione: (2025)
di: Cattan, Arie, et al.
Pubblicazione: (2025)
Latent Reasoning with Supervised Thinking States
di: Amos, Ido, et al.
Pubblicazione: (2026)
di: Amos, Ido, et al.
Pubblicazione: (2026)
A Novel Computational and Modeling Foundation for Automatic Coherence Assessment
di: Maimon, Aviya, et al.
Pubblicazione: (2023)
di: Maimon, Aviya, et al.
Pubblicazione: (2023)
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
di: Goldman, Omer, et al.
Pubblicazione: (2024)
di: Goldman, Omer, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
di: Basmov, Victoria, et al.
Pubblicazione: (2023)
di: Basmov, Victoria, et al.
Pubblicazione: (2023)
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2024)
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
di: Ghandeharioun, Asma, et al.
Pubblicazione: (2024)
di: Ghandeharioun, Asma, et al.
Pubblicazione: (2024)
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew
di: Paz-Argaman, Tzuf, et al.
Pubblicazione: (2024)
di: Paz-Argaman, Tzuf, et al.
Pubblicazione: (2024)
Breaking the Language Barrier: Can Direct Inference Outperform Pre-Translation in Multilingual LLM Applications?
di: Intrator, Yotam, et al.
Pubblicazione: (2024)
di: Intrator, Yotam, et al.
Pubblicazione: (2024)
HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
di: Cohen, Amir DN, et al.
Pubblicazione: (2025)
di: Cohen, Amir DN, et al.
Pubblicazione: (2025)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2025)
Into the Unknown: Generating Geospatial Descriptions for New Environments
di: Paz-Argaman, Tzuf, et al.
Pubblicazione: (2024)
di: Paz-Argaman, Tzuf, et al.
Pubblicazione: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
Correlated Errors in Large Language Models
di: Kim, Elliot, et al.
Pubblicazione: (2025)
di: Kim, Elliot, et al.
Pubblicazione: (2025)
Lossless Token Sequence Compression via Meta-Tokens
di: Harvill, John, et al.
Pubblicazione: (2025)
di: Harvill, John, et al.
Pubblicazione: (2025)
Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image Models
di: Ventura, Mor, et al.
Pubblicazione: (2023)
di: Ventura, Mor, et al.
Pubblicazione: (2023)
Identifying User Goals from UI Trajectories
di: Berkovitch, Omri, et al.
Pubblicazione: (2024)
di: Berkovitch, Omri, et al.
Pubblicazione: (2024)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
di: Bowyer, Sam, et al.
Pubblicazione: (2026)
di: Bowyer, Sam, et al.
Pubblicazione: (2026)
MRL Parsing Without Tears: The Case of Hebrew
di: Shmidman, Shaltiel, et al.
Pubblicazione: (2024)
di: Shmidman, Shaltiel, et al.
Pubblicazione: (2024)
Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding
di: Zhang, Zhongjian, et al.
Pubblicazione: (2026)
di: Zhang, Zhongjian, et al.
Pubblicazione: (2026)
UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation
di: Guo, Shuhan, et al.
Pubblicazione: (2024)
di: Guo, Shuhan, et al.
Pubblicazione: (2024)
Training Large Language Models to Predict Clinical Events
di: Turtel, Benjamin, et al.
Pubblicazione: (2026)
di: Turtel, Benjamin, et al.
Pubblicazione: (2026)
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
di: Wolfson, Tomer, et al.
Pubblicazione: (2025)
di: Wolfson, Tomer, et al.
Pubblicazione: (2025)
Evaluation of Large Language Models via Coupled Token Generation
di: Benz, Nina Corvelo, et al.
Pubblicazione: (2025)
di: Benz, Nina Corvelo, et al.
Pubblicazione: (2025)
SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
di: Tănase, Andrei-Valentin, et al.
Pubblicazione: (2025)
di: Tănase, Andrei-Valentin, et al.
Pubblicazione: (2025)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
di: Goel, Vatsal, et al.
Pubblicazione: (2026)
di: Goel, Vatsal, et al.
Pubblicazione: (2026)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
di: Mittu, Fazal, et al.
Pubblicazione: (2024)
di: Mittu, Fazal, et al.
Pubblicazione: (2024)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
di: Su, DiJia, et al.
Pubblicazione: (2025)
di: Su, DiJia, et al.
Pubblicazione: (2025)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
di: Cao, Qi, et al.
Pubblicazione: (2026)
di: Cao, Qi, et al.
Pubblicazione: (2026)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
di: Lu, Liming, et al.
Pubblicazione: (2026)
di: Lu, Liming, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
di: Shaham, Uri, et al.
Pubblicazione: (2024) -
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
di: Mor-Lan, Guy, et al.
Pubblicazione: (2026) -
ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer
di: Goldman, Omer, et al.
Pubblicazione: (2025) -
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
di: Greenfeld, Refael Shaked, et al.
Pubblicazione: (2026) -
A Truly Joint Neural Architecture for Segmentation and Parsing
di: Levi, Danit Yshaayahu, et al.
Pubblicazione: (2024)