X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Sreenivas, Sharath Turuvekere, Hanasoge, Adithyakrishna Venkatesh, Yang, Mingyu, Taghibakhshi, Ali, Muralidharan, Saurav, Aithal, Ashwath, Molchanov, Pavlo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Compact Language Models via Pruning and Knowledge Distillation
di: Muralidharan, Saurav, et al.
Pubblicazione: (2024)
di: Muralidharan, Saurav, et al.
Pubblicazione: (2024)
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024)
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024)
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2026)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2026)
Flextron: Many-in-One Flexible Large Language Model
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
di: Fang, Gongfan, et al.
Pubblicazione: (2024)
di: Fang, Gongfan, et al.
Pubblicazione: (2024)
Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
di: Venkatesh, Sohan
Pubblicazione: (2026)
di: Venkatesh, Sohan
Pubblicazione: (2026)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
di: Phan, Buu, et al.
Pubblicazione: (2025)
di: Phan, Buu, et al.
Pubblicazione: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
di: Shen, Gerald, et al.
Pubblicazione: (2024)
di: Shen, Gerald, et al.
Pubblicazione: (2024)
Multi-Token Prediction via Self-Distillation
di: Kirchenbauer, John, et al.
Pubblicazione: (2026)
di: Kirchenbauer, John, et al.
Pubblicazione: (2026)
Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping
di: Chen, Yijie, et al.
Pubblicazione: (2025)
di: Chen, Yijie, et al.
Pubblicazione: (2025)
Self-Distillation for Multi-Token Prediction
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
di: Liu, Shih-Yang, et al.
Pubblicazione: (2025)
di: Liu, Shih-Yang, et al.
Pubblicazione: (2025)
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
di: Zhang, Hao, et al.
Pubblicazione: (2026)
di: Zhang, Hao, et al.
Pubblicazione: (2026)
CTPD: Cross Tokenizer Preference Distillation
di: Nguyen, Truong, et al.
Pubblicazione: (2026)
di: Nguyen, Truong, et al.
Pubblicazione: (2026)
Guiding Language Model Reasoning with Planning Tokens
di: Wang, Xinyi, et al.
Pubblicazione: (2023)
di: Wang, Xinyi, et al.
Pubblicazione: (2023)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
di: Kim, Minsang, et al.
Pubblicazione: (2026)
di: Kim, Minsang, et al.
Pubblicazione: (2026)
Persuasion Tokens for Editing Factual Knowledge in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2026)
di: Youssef, Paul, et al.
Pubblicazione: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
Mitigating Heterogeneous Token Overfitting in LLM Knowledge Editing
di: Liu, Tianci, et al.
Pubblicazione: (2025)
di: Liu, Tianci, et al.
Pubblicazione: (2025)
X-VILA: Cross-Modality Alignment for Large Language Model
di: Ye, Hanrong, et al.
Pubblicazione: (2024)
di: Ye, Hanrong, et al.
Pubblicazione: (2024)
Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models
di: Cui, Xiao, et al.
Pubblicazione: (2024)
di: Cui, Xiao, et al.
Pubblicazione: (2024)
LLM-Oriented Token-Adaptive Knowledge Distillation
di: Xie, Xurong, et al.
Pubblicazione: (2025)
di: Xie, Xurong, et al.
Pubblicazione: (2025)
TokenButler: Token Importance is Predictable
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
di: Boizard, Nicolas, et al.
Pubblicazione: (2024)
di: Boizard, Nicolas, et al.
Pubblicazione: (2024)
TokenShapley: Token Level Context Attribution with Shapley Value
di: Xiao, Yingtai, et al.
Pubblicazione: (2025)
di: Xiao, Yingtai, et al.
Pubblicazione: (2025)
Towards Token-Level Text Anomaly Detection
di: Cao, Yang, et al.
Pubblicazione: (2026)
di: Cao, Yang, et al.
Pubblicazione: (2026)
DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation
di: Vu, Duc Trung, et al.
Pubblicazione: (2026)
di: Vu, Duc Trung, et al.
Pubblicazione: (2026)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
di: Antoniak, Szymon, et al.
Pubblicazione: (2023)
di: Antoniak, Szymon, et al.
Pubblicazione: (2023)
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
di: Wu, Ning, et al.
Pubblicazione: (2026)
di: Wu, Ning, et al.
Pubblicazione: (2026)
Lossless Token Sequence Compression via Meta-Tokens
di: Harvill, John, et al.
Pubblicazione: (2025)
di: Harvill, John, et al.
Pubblicazione: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
di: Jo, Dongwon, et al.
Pubblicazione: (2026)
di: Jo, Dongwon, et al.
Pubblicazione: (2026)
Universal Cross-Tokenizer Distillation via Approximate Likelihood Matching
di: Minixhofer, Benjamin, et al.
Pubblicazione: (2025)
di: Minixhofer, Benjamin, et al.
Pubblicazione: (2025)
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
di: Guo, Yiju, et al.
Pubblicazione: (2025)
di: Guo, Yiju, et al.
Pubblicazione: (2025)
Adversarial Tokenization
di: Geh, Renato Lui, et al.
Pubblicazione: (2025)
di: Geh, Renato Lui, et al.
Pubblicazione: (2025)
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
di: Li, Shiwei, et al.
Pubblicazione: (2025)
di: Li, Shiwei, et al.
Pubblicazione: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Compact Language Models via Pruning and Knowledge Distillation
di: Muralidharan, Saurav, et al.
Pubblicazione: (2024) -
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025) -
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025) -
LLM Pruning and Distillation in Practice: The Minitron Approach
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024) -
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2026)