Token Distillation: Attention-aware Input Embeddings For New Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Dobler, Konstantin, Elliott, Desmond, de Melo, Gerard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language Adaptation on a Tight Academic Compute Budget: Tokenizer Swapping Works and Pure bfloat16 Is Enough
by: Dobler, Konstantin, et al.
Published: (2024)
by: Dobler, Konstantin, et al.
Published: (2024)
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
by: Cohen, Roi, et al.
Published: (2024)
by: Cohen, Roi, et al.
Published: (2024)
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
by: Dobler, Konstantin, et al.
Published: (2023)
by: Dobler, Konstantin, et al.
Published: (2023)
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
Measuring Intrinsic Dimension of Token Embeddings
by: Kataiwa, Takuya, et al.
Published: (2025)
by: Kataiwa, Takuya, et al.
Published: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
Multi-Token Prediction via Self-Distillation
by: Kirchenbauer, John, et al.
Published: (2026)
by: Kirchenbauer, John, et al.
Published: (2026)
Tracking Universal Features Through Fine-Tuning and Model Merging
by: Horn, Niels, et al.
Published: (2024)
by: Horn, Niels, et al.
Published: (2024)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization
by: Deekeswar, Harshavardhanan
Published: (2026)
by: Deekeswar, Harshavardhanan
Published: (2026)
Self-Distillation for Multi-Token Prediction
by: Zhao, Guoliang, et al.
Published: (2026)
by: Zhao, Guoliang, et al.
Published: (2026)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
by: Phan, Buu, et al.
Published: (2025)
by: Phan, Buu, et al.
Published: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
by: Xu, Ceyu, et al.
Published: (2026)
by: Xu, Ceyu, et al.
Published: (2026)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025)
by: He, Mutian, et al.
Published: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
by: Deng, Difan, et al.
Published: (2026)
by: Deng, Difan, et al.
Published: (2026)
Understanding Token Probability Encoding in Output Embeddings
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
by: Li, Shiwei, et al.
Published: (2025)
by: Li, Shiwei, et al.
Published: (2025)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
by: Zhang, Songming, et al.
Published: (2025)
by: Zhang, Songming, et al.
Published: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
by: Mihaila, George
Published: (2026)
by: Mihaila, George
Published: (2026)
Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence
by: Işık, İlker, et al.
Published: (2024)
by: Işık, İlker, et al.
Published: (2024)
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
by: Ding, Xueying, et al.
Published: (2025)
by: Ding, Xueying, et al.
Published: (2025)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
TokenShapley: Token Level Context Attribution with Shapley Value
by: Xiao, Yingtai, et al.
Published: (2025)
by: Xiao, Yingtai, et al.
Published: (2025)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
by: Addanki, Raghav, et al.
Published: (2023)
by: Addanki, Raghav, et al.
Published: (2023)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
by: Kumar, Gokul Karthik, et al.
Published: (2026)
by: Kumar, Gokul Karthik, et al.
Published: (2026)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
by: Liu, Yichen, et al.
Published: (2022)
by: Liu, Yichen, et al.
Published: (2022)
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
by: Wu, Ning, et al.
Published: (2026)
by: Wu, Ning, et al.
Published: (2026)
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
by: Song, Guanghui, et al.
Published: (2025)
by: Song, Guanghui, et al.
Published: (2025)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
by: Goel, Raghavv, et al.
Published: (2025)
by: Goel, Raghavv, et al.
Published: (2025)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
by: Jo, Hyun-rae, et al.
Published: (2024)
by: Jo, Hyun-rae, et al.
Published: (2024)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Unsupervised Morphological Tree Tokenizer
by: Zhu, Qingyang, et al.
Published: (2024)
by: Zhu, Qingyang, et al.
Published: (2024)
Similar Items
-
Language Adaptation on a Tight Academic Compute Budget: Tokenizer Swapping Works and Pure bfloat16 Is Enough
by: Dobler, Konstantin, et al.
Published: (2024) -
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
by: Cohen, Roi, et al.
Published: (2024) -
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
by: Dobler, Konstantin, et al.
Published: (2023) -
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025) -
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)