Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Boizard, Nicolas, Haddad, Kevin El, Hudelot, Céline, Colombo, Pierre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs
von: Boizard, Nicolas, et al.
Veröffentlicht: (2026)
von: Boizard, Nicolas, et al.
Veröffentlicht: (2026)
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2026)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2026)
Should We Still Pretrain Encoders with Masked Language Modeling?
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025)
DistillLens: Symmetric Knowledge Distillation Through Logit Lens
von: Dhakal, Manish, et al.
Veröffentlicht: (2026)
von: Dhakal, Manish, et al.
Veröffentlicht: (2026)
Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2024)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2024)
Universal Cross-Tokenizer Distillation via Approximate Likelihood Matching
von: Minixhofer, Benjamin, et al.
Veröffentlicht: (2025)
von: Minixhofer, Benjamin, et al.
Veröffentlicht: (2025)
CTPD: Cross Tokenizer Preference Distillation
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
von: Li, Minchong, et al.
Veröffentlicht: (2024)
von: Li, Minchong, et al.
Veröffentlicht: (2024)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
von: Anshumann, et al.
Veröffentlicht: (2025)
von: Anshumann, et al.
Veröffentlicht: (2025)
LoCa: Logit Calibration for Knowledge Distillation
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping
von: Chen, Yijie, et al.
Veröffentlicht: (2025)
von: Chen, Yijie, et al.
Veröffentlicht: (2025)
Cross-Tokenizer LLM Distillation through a Byte-Level Interface
von: Singh, Avyav Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Avyav Kumar, et al.
Veröffentlicht: (2026)
SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation
von: Sun, Jie, et al.
Veröffentlicht: (2026)
von: Sun, Jie, et al.
Veröffentlicht: (2026)
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2024)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2024)
EuroBERT: Scaling Multilingual Encoders for European Languages
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
von: Phan, Buu, et al.
Veröffentlicht: (2025)
von: Phan, Buu, et al.
Veröffentlicht: (2025)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
von: Aswal, Darpan, et al.
Veröffentlicht: (2025)
von: Aswal, Darpan, et al.
Veröffentlicht: (2025)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model Fusion
von: Wang, Yuanyi, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyi, et al.
Veröffentlicht: (2025)
Enhancing Knowledge Distillation for LLMs with Response-Priming Prompting
von: Goyal, Vijay, et al.
Veröffentlicht: (2024)
von: Goyal, Vijay, et al.
Veröffentlicht: (2024)
UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models
von: Dong, Yijiang River, et al.
Veröffentlicht: (2024)
von: Dong, Yijiang River, et al.
Veröffentlicht: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
ColPali: Efficient Document Retrieval with Vision Language Models
von: Faysse, Manuel, et al.
Veröffentlicht: (2024)
von: Faysse, Manuel, et al.
Veröffentlicht: (2024)
DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation
von: Vu, Duc Trung, et al.
Veröffentlicht: (2026)
von: Vu, Duc Trung, et al.
Veröffentlicht: (2026)
Distilling Token-Trained Models into Byte-Level Models
von: Bao, Zishuo, et al.
Veröffentlicht: (2026)
von: Bao, Zishuo, et al.
Veröffentlicht: (2026)
Token Distillation: Attention-aware Input Embeddings For New Tokens
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
von: Du, Dayou, et al.
Veröffentlicht: (2024)
von: Du, Dayou, et al.
Veröffentlicht: (2024)
LLM-Oriented Token-Adaptive Knowledge Distillation
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
Hybrid Policy Distillation for LLMs
von: Zhu, Wenhong, et al.
Veröffentlicht: (2026)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2026)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
Translate-Distill: Learning Cross-Language Dense Retrieval by Translation and Distillation
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
von: Li, Jie, et al.
Veröffentlicht: (2024)
von: Li, Jie, et al.
Veröffentlicht: (2024)
Distill-C: Enhanced NL2SQL via Distilled Customization with LLMs
von: Hoang, Cong Duy Vu, et al.
Veröffentlicht: (2025)
von: Hoang, Cong Duy Vu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
von: Boizard, Nicolas, et al.
Veröffentlicht: (2025) -
BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs
von: Boizard, Nicolas, et al.
Veröffentlicht: (2026) -
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2026) -
Should We Still Pretrain Encoders with Masked Language Modeling?
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2025) -
DistillLens: Symmetric Knowledge Distillation Through Logit Lens
von: Dhakal, Manish, et al.
Veröffentlicht: (2026)