Saved in:
| Main Authors: | R V, Kavin, Goyal, Pawan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.17737 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation
by: Belikova, Julia, et al.
Published: (2026)
by: Belikova, Julia, et al.
Published: (2026)
Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation
by: Valois, Pedro H. V., et al.
Published: (2024)
by: Valois, Pedro H. V., et al.
Published: (2024)
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
by: Zhang, Christine, et al.
Published: (2026)
by: Zhang, Christine, et al.
Published: (2026)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025)
by: Nakkiran, Preetum, et al.
Published: (2025)
Vector Arithmetic in Concept and Token Subspaces
by: Feucht, Sheridan, et al.
Published: (2025)
by: Feucht, Sheridan, et al.
Published: (2025)
xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
by: Cheng, Xin, et al.
Published: (2024)
by: Cheng, Xin, et al.
Published: (2024)
Lossless Token Sequence Compression via Meta-Tokens
by: Harvill, John, et al.
Published: (2025)
by: Harvill, John, et al.
Published: (2025)
Beyond Text Compression: Evaluating Tokenizers Across Scales
by: Lotz, Jonas F., et al.
Published: (2025)
by: Lotz, Jonas F., et al.
Published: (2025)
Learning to Compress Prompts with Gist Tokens
by: Mu, Jesse, et al.
Published: (2023)
by: Mu, Jesse, et al.
Published: (2023)
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
by: Zhao, Xinye, et al.
Published: (2025)
by: Zhao, Xinye, et al.
Published: (2025)
Multi-word Tokenization for Sequence Compression
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Beyond Tokens: Concept-Level Training Objectives for LLMs
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
by: Lin, Xiaolin, et al.
Published: (2025)
by: Lin, Xiaolin, et al.
Published: (2025)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
by: Zhao, Yize, et al.
Published: (2025)
by: Zhao, Yize, et al.
Published: (2025)
From Tokens to Concepts: Leveraging SAE for SPLADE
by: Zong, Yuxuan, et al.
Published: (2026)
by: Zong, Yuxuan, et al.
Published: (2026)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
by: Thakur, Aamod, et al.
Published: (2025)
by: Thakur, Aamod, et al.
Published: (2025)
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
by: Zhang, Jiebin, et al.
Published: (2024)
by: Zhang, Jiebin, et al.
Published: (2024)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
by: Song, Yuhan, et al.
Published: (2025)
by: Song, Yuhan, et al.
Published: (2025)
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
by: An, Ruize, et al.
Published: (2025)
by: An, Ruize, et al.
Published: (2025)
Semantic Tokens in Retrieval Augmented Generation
by: Suro, Joel
Published: (2024)
by: Suro, Joel
Published: (2024)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)
by: Schmidt, Craig W., et al.
Published: (2024)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
by: Zhu, Mingcheng, et al.
Published: (2026)
by: Zhu, Mingcheng, et al.
Published: (2026)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)
by: Ovalle, Anaelia, et al.
Published: (2023)
Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
by: Bochkov, A.
Published: (2025)
by: Bochkov, A.
Published: (2025)
Error-Aware Curriculum Learning for Biomedical Relation Classification
by: Chakraborty, Sinchani, et al.
Published: (2025)
by: Chakraborty, Sinchani, et al.
Published: (2025)
Token Sequence Compression for Efficient Multimodal Computing
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
Intent Detection and Entity Extraction from BioMedical Literature
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
Vision-centric Token Compression in Large Language Model
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
Efficient Whole Slide Pathology VQA via Token Compression
by: Lyu, Weimin, et al.
Published: (2025)
by: Lyu, Weimin, et al.
Published: (2025)
Hypernym Mercury: Token Optimization Through Semantic Field Constriction And Reconstruction From Hypernyms. A New Text Compression Method
by: Forrester, Chris, et al.
Published: (2025)
by: Forrester, Chris, et al.
Published: (2025)
Contextual Reinforcement in Multimodal Token Compression for Large Language Models
by: Piero, Naderdel, et al.
Published: (2025)
by: Piero, Naderdel, et al.
Published: (2025)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
by: Li, Zilong
Published: (2024)
by: Li, Zilong
Published: (2024)
Beyond Literal Token Overlap: Token Alignability for Multilinguality
by: Hämmerl, Katharina, et al.
Published: (2025)
by: Hämmerl, Katharina, et al.
Published: (2025)
Problematic Tokens: Tokenizer Bias in Large Language Models
by: Yang, Jin, et al.
Published: (2024)
by: Yang, Jin, et al.
Published: (2024)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
by: Walker, Nicholas
Published: (2024)
by: Walker, Nicholas
Published: (2024)
Similar Items
-
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
by: Goyal, Agam, et al.
Published: (2025) -
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026) -
Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation
by: Belikova, Julia, et al.
Published: (2026) -
Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text Generation
by: Valois, Pedro H. V., et al.
Published: (2024) -
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
by: Zhang, Christine, et al.
Published: (2026)