Beyond Text Compression: Evaluating Tokenizers Across Scales
Fuente:
arXiv
Salvato in:
| Autori principali: | Lotz, Jonas F., Lopes, António V., Peitz, Stephan, Setiawan, Hendra, Emili, Leonardo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Overcoming Vocabulary Constraints with Pixel-level Fallback
di: Lotz, Jonas F., et al.
Pubblicazione: (2025)
di: Lotz, Jonas F., et al.
Pubblicazione: (2025)
Accurate Knowledge Distillation with n-best Reranking
di: Setiawan, Hendra
Pubblicazione: (2023)
di: Setiawan, Hendra
Pubblicazione: (2023)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
di: Goldman, Omer, et al.
Pubblicazione: (2024)
di: Goldman, Omer, et al.
Pubblicazione: (2024)
Multi-word Tokenization for Sequence Compression
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
Text Generation Beyond Discrete Token Sampling
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches
di: Pratapa, Adithya, et al.
Pubblicazione: (2025)
di: Pratapa, Adithya, et al.
Pubblicazione: (2025)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
di: Kaplan, Guy, et al.
Pubblicazione: (2025)
di: Kaplan, Guy, et al.
Pubblicazione: (2025)
The Role of Data Curation in Image Captioning
di: Li, Wenyan, et al.
Pubblicazione: (2023)
di: Li, Wenyan, et al.
Pubblicazione: (2023)
Scaling Optimal LR Across Token Horizons
di: Bjorck, Johan, et al.
Pubblicazione: (2024)
di: Bjorck, Johan, et al.
Pubblicazione: (2024)
Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
di: R V, Kavin, et al.
Pubblicazione: (2025)
di: R V, Kavin, et al.
Pubblicazione: (2025)
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
di: Perlitz, Yotam, et al.
Pubblicazione: (2026)
di: Perlitz, Yotam, et al.
Pubblicazione: (2026)
HRM-Text: Efficient Pretraining Beyond Scaling
di: Wang, Guan, et al.
Pubblicazione: (2026)
di: Wang, Guan, et al.
Pubblicazione: (2026)
Multilingual Pretraining for Pixel Language Models
di: Kesen, Ilker, et al.
Pubblicazione: (2025)
di: Kesen, Ilker, et al.
Pubblicazione: (2025)
An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification
di: Rusli, Andre, et al.
Pubblicazione: (2024)
di: Rusli, Andre, et al.
Pubblicazione: (2024)
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Learning to Compress Prompts with Gist Tokens
di: Mu, Jesse, et al.
Pubblicazione: (2023)
di: Mu, Jesse, et al.
Pubblicazione: (2023)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
di: Larionov, Daniil, et al.
Pubblicazione: (2025)
di: Larionov, Daniil, et al.
Pubblicazione: (2025)
Beyond Literal Token Overlap: Token Alignability for Multilinguality
di: Hämmerl, Katharina, et al.
Pubblicazione: (2025)
di: Hämmerl, Katharina, et al.
Pubblicazione: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
di: Mao, Yu, et al.
Pubblicazione: (2025)
di: Mao, Yu, et al.
Pubblicazione: (2025)
$\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
di: Zhang, Xinrong, et al.
Pubblicazione: (2024)
di: Zhang, Xinrong, et al.
Pubblicazione: (2024)
Frequency-Ordered Tokenization for Better Text Compression
di: Kalcher, Maximilian
Pubblicazione: (2026)
di: Kalcher, Maximilian
Pubblicazione: (2026)
Beyond the Needle's Illusion: Decoupled Evaluation of Evidence Access and Use under Semantic Interference at 326M-Token Scale
di: Lin, Tianwei, et al.
Pubblicazione: (2026)
di: Lin, Tianwei, et al.
Pubblicazione: (2026)
Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages
di: Tamang, S., et al.
Pubblicazione: (2024)
di: Tamang, S., et al.
Pubblicazione: (2024)
Evaluating LLMs and Pre-trained Models for Text Summarization Across Diverse Datasets
di: Rehman, Tohida, et al.
Pubblicazione: (2025)
di: Rehman, Tohida, et al.
Pubblicazione: (2025)
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
di: Ying, Shuangshuang, et al.
Pubblicazione: (2025)
di: Ying, Shuangshuang, et al.
Pubblicazione: (2025)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
di: van Noord, Rik, et al.
Pubblicazione: (2024)
di: van Noord, Rik, et al.
Pubblicazione: (2024)
Text2Cypher Across Languages: Evaluating and Finetuning LLMs
di: Ozsoy, Makbule Gulcin, et al.
Pubblicazione: (2025)
di: Ozsoy, Makbule Gulcin, et al.
Pubblicazione: (2025)
Hypernym Mercury: Token Optimization Through Semantic Field Constriction And Reconstruction From Hypernyms. A New Text Compression Method
di: Forrester, Chris, et al.
Pubblicazione: (2025)
di: Forrester, Chris, et al.
Pubblicazione: (2025)
Matina: A Large-Scale 73B Token Persian Text Corpus
di: Hosseinbeigi, Sara Bourbour, et al.
Pubblicazione: (2025)
di: Hosseinbeigi, Sara Bourbour, et al.
Pubblicazione: (2025)
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
di: An, Ruize, et al.
Pubblicazione: (2025)
di: An, Ruize, et al.
Pubblicazione: (2025)
Glyph: Scaling Context Windows via Visual-Text Compression
di: Cheng, Jiale, et al.
Pubblicazione: (2025)
di: Cheng, Jiale, et al.
Pubblicazione: (2025)
Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
di: Benito-Rodriguez, Éloïse, et al.
Pubblicazione: (2025)
di: Benito-Rodriguez, Éloïse, et al.
Pubblicazione: (2025)
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
di: Zhang, Jiebin, et al.
Pubblicazione: (2024)
di: Zhang, Jiebin, et al.
Pubblicazione: (2024)
Lossless Token Sequence Compression via Meta-Tokens
di: Harvill, John, et al.
Pubblicazione: (2025)
di: Harvill, John, et al.
Pubblicazione: (2025)
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
di: Zheng, Hao, et al.
Pubblicazione: (2025)
di: Zheng, Hao, et al.
Pubblicazione: (2025)
DiffScore: Text Evaluation Beyond Autoregressive Likelihood
di: Lai, Wen, et al.
Pubblicazione: (2026)
di: Lai, Wen, et al.
Pubblicazione: (2026)
Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation
di: Belikova, Julia, et al.
Pubblicazione: (2026)
di: Belikova, Julia, et al.
Pubblicazione: (2026)
Text Compression for Efficient Language Generation
di: Gu, David, et al.
Pubblicazione: (2025)
di: Gu, David, et al.
Pubblicazione: (2025)
Building Consumer Loyalty: Understanding E-Satisfaction in Fast-Fashion Purchases on E-Commerce Platform
di: Wijaya, Hendra, et al.
Pubblicazione: (2025)
di: Wijaya, Hendra, et al.
Pubblicazione: (2025)
Beyond Tokens: Concept-Level Training Objectives for LLMs
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Overcoming Vocabulary Constraints with Pixel-level Fallback
di: Lotz, Jonas F., et al.
Pubblicazione: (2025) -
Accurate Knowledge Distillation with n-best Reranking
di: Setiawan, Hendra
Pubblicazione: (2023) -
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
di: Goldman, Omer, et al.
Pubblicazione: (2024) -
Multi-word Tokenization for Sequence Compression
di: Gee, Leonidas, et al.
Pubblicazione: (2024) -
Text Generation Beyond Discrete Token Sampling
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)