Hypernym Mercury: Token Optimization Through Semantic Field Constriction And Reconstruction From Hypernyms. A New Text Compression Method
Fuente:
arXiv
Salvato in:
| Autori principali: | Forrester, Chris, Sulea, Octavia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Prompt-Based Methods for Zero-Shot Hypernym Prediction with Large Language Models
di: Tikhomirov, Mikhail, et al.
Pubblicazione: (2024)
di: Tikhomirov, Mikhail, et al.
Pubblicazione: (2024)
SHADE: Semantic Hypernym Annotator for Domain-specific Entities -- DnD Domain Use Case
di: Peiris, Akila, et al.
Pubblicazione: (2024)
di: Peiris, Akila, et al.
Pubblicazione: (2024)
Inferring Adjective Hypernyms with Language Models to Increase the Connectivity of Open English Wordnet
di: Augello, Lorenzo, et al.
Pubblicazione: (2025)
di: Augello, Lorenzo, et al.
Pubblicazione: (2025)
HyperBox: A Supervised Approach for Hypernym Discovery using Box Embeddings
di: Parmar, Maulik, et al.
Pubblicazione: (2022)
di: Parmar, Maulik, et al.
Pubblicazione: (2022)
Hypernym Bias: Unraveling Deep Classifier Training Dynamics through the Lens of Class Hierarchy
di: Malashin, Roman, et al.
Pubblicazione: (2025)
di: Malashin, Roman, et al.
Pubblicazione: (2025)
On the Semantic and Syntactic Information Encoded in Proto-Tokens for One-Step Text Reconstruction
di: Bondarenko, Ivan, et al.
Pubblicazione: (2026)
di: Bondarenko, Ivan, et al.
Pubblicazione: (2026)
Beyond Text Compression: Evaluating Tokenizers Across Scales
di: Lotz, Jonas F., et al.
Pubblicazione: (2025)
di: Lotz, Jonas F., et al.
Pubblicazione: (2025)
Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
di: R V, Kavin, et al.
Pubblicazione: (2025)
di: R V, Kavin, et al.
Pubblicazione: (2025)
Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction
di: Zou, Yuchun, et al.
Pubblicazione: (2026)
di: Zou, Yuchun, et al.
Pubblicazione: (2026)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
di: Yuan, Aomufei, et al.
Pubblicazione: (2025)
di: Yuan, Aomufei, et al.
Pubblicazione: (2025)
See the Text: From Tokenization to Visual Reading
di: Xing, Ling, et al.
Pubblicazione: (2025)
di: Xing, Ling, et al.
Pubblicazione: (2025)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
di: Uzan, Omri, et al.
Pubblicazione: (2024)
di: Uzan, Omri, et al.
Pubblicazione: (2024)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
di: Zhu, Mingcheng, et al.
Pubblicazione: (2026)
di: Zhu, Mingcheng, et al.
Pubblicazione: (2026)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
di: Lin, Xiaolin, et al.
Pubblicazione: (2025)
di: Lin, Xiaolin, et al.
Pubblicazione: (2025)
Learning to Compress Prompts with Gist Tokens
di: Mu, Jesse, et al.
Pubblicazione: (2023)
di: Mu, Jesse, et al.
Pubblicazione: (2023)
ACT-MNMT Auto-Constriction Turning for Multilingual Neural Machine Translation
di: Dai, Shaojie, et al.
Pubblicazione: (2024)
di: Dai, Shaojie, et al.
Pubblicazione: (2024)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
di: Goldman, Omer, et al.
Pubblicazione: (2024)
di: Goldman, Omer, et al.
Pubblicazione: (2024)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
di: Mao, Yu, et al.
Pubblicazione: (2025)
di: Mao, Yu, et al.
Pubblicazione: (2025)
zip2zip: Inference-Time Adaptive Tokenization via Online Compression
di: Geng, Saibo, et al.
Pubblicazione: (2025)
di: Geng, Saibo, et al.
Pubblicazione: (2025)
Frequency-Ordered Tokenization for Better Text Compression
di: Kalcher, Maximilian
Pubblicazione: (2026)
di: Kalcher, Maximilian
Pubblicazione: (2026)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
di: Zhao, Yize, et al.
Pubblicazione: (2025)
di: Zhao, Yize, et al.
Pubblicazione: (2025)
LLM-Augmented Semantic Steering of Text Embedding Projection Spaces
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction
di: Yang, Xiaoli, et al.
Pubblicazione: (2026)
di: Yang, Xiaoli, et al.
Pubblicazione: (2026)
Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary Adaptation
di: Moroni, Luca, et al.
Pubblicazione: (2025)
di: Moroni, Luca, et al.
Pubblicazione: (2025)
Faster Superword Tokenization
di: Schmidt, Craig W., et al.
Pubblicazione: (2026)
di: Schmidt, Craig W., et al.
Pubblicazione: (2026)
From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution
di: Chizhov, Pavel, et al.
Pubblicazione: (2026)
di: Chizhov, Pavel, et al.
Pubblicazione: (2026)
Morphologically-Informed Tokenizers for Languages with Non-Concatenative Morphology: A case study of Yoloxóchtil Mixtec ASR
di: Crawford, Chris
Pubblicazione: (2025)
di: Crawford, Chris
Pubblicazione: (2025)
SemanticZip: A Pilot Framework for Lossy Text Compression with LLMs as Semantic Decompressors
di: Trukhina, Natalia, et al.
Pubblicazione: (2026)
di: Trukhina, Natalia, et al.
Pubblicazione: (2026)
Multi-word Tokenization for Sequence Compression
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
di: An, Ruize, et al.
Pubblicazione: (2025)
di: An, Ruize, et al.
Pubblicazione: (2025)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
di: Shani, Chen, et al.
Pubblicazione: (2025)
di: Shani, Chen, et al.
Pubblicazione: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
di: Song, Yuhan, et al.
Pubblicazione: (2025)
di: Song, Yuhan, et al.
Pubblicazione: (2025)
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
di: Zhang, Jiebin, et al.
Pubblicazione: (2024)
di: Zhang, Jiebin, et al.
Pubblicazione: (2024)
Lossless Token Sequence Compression via Meta-Tokens
di: Harvill, John, et al.
Pubblicazione: (2025)
di: Harvill, John, et al.
Pubblicazione: (2025)
Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation
di: Belikova, Julia, et al.
Pubblicazione: (2026)
di: Belikova, Julia, et al.
Pubblicazione: (2026)
Text Compression for Efficient Language Generation
di: Gu, David, et al.
Pubblicazione: (2025)
di: Gu, David, et al.
Pubblicazione: (2025)
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
di: Nie, Zhijie, et al.
Pubblicazione: (2024)
di: Nie, Zhijie, et al.
Pubblicazione: (2024)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
di: Manakul, Potsawee, et al.
Pubblicazione: (2026)
di: Manakul, Potsawee, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Exploring Prompt-Based Methods for Zero-Shot Hypernym Prediction with Large Language Models
di: Tikhomirov, Mikhail, et al.
Pubblicazione: (2024) -
SHADE: Semantic Hypernym Annotator for Domain-specific Entities -- DnD Domain Use Case
di: Peiris, Akila, et al.
Pubblicazione: (2024) -
Inferring Adjective Hypernyms with Language Models to Increase the Connectivity of Open English Wordnet
di: Augello, Lorenzo, et al.
Pubblicazione: (2025) -
HyperBox: A Supervised Approach for Hypernym Discovery using Box Embeddings
di: Parmar, Maulik, et al.
Pubblicazione: (2022) -
Hypernym Bias: Unraveling Deep Classifier Training Dynamics through the Lens of Class Hierarchy
di: Malashin, Roman, et al.
Pubblicazione: (2025)