Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phan, Buu, Khisti, Ashish, Ullrich, Karen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling
von: Phan, Buu, et al.
Veröffentlicht: (2025)
von: Phan, Buu, et al.
Veröffentlicht: (2025)
List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
Multi-Marginal Couplings for Metropolis-Hastings
von: Phan, Buu, et al.
Veröffentlicht: (2026)
von: Phan, Buu, et al.
Veröffentlicht: (2026)
Importance Matching Lemma for Lossy Compression with Side Information
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
One-Shot Broadcast Joint Source-Channel Coding with Codebook Diversity
von: Rowan, Joseph, et al.
Veröffentlicht: (2026)
von: Rowan, Joseph, et al.
Veröffentlicht: (2026)
Token Distillation: Attention-aware Input Embeddings For New Tokens
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
On Self-Adaptive Perception Loss Function for Sequential Lossy Compression
von: Salehkalaibar, Sadaf, et al.
Veröffentlicht: (2025)
von: Salehkalaibar, Sadaf, et al.
Veröffentlicht: (2025)
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models
von: Huo, Mingjia, et al.
Veröffentlicht: (2024)
von: Huo, Mingjia, et al.
Veröffentlicht: (2024)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
Language Modeling with Learned Meta-Tokens
von: Shah, Alok N., et al.
Veröffentlicht: (2025)
von: Shah, Alok N., et al.
Veröffentlicht: (2025)
Parallel Token Prediction for Language Models
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
End-To-End Causal Effect Estimation from Unstructured Natural Language Data
von: Dhawan, Nikita, et al.
Veröffentlicht: (2024)
von: Dhawan, Nikita, et al.
Veröffentlicht: (2024)
SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
von: Tănase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
von: Tănase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
Multi-Draft Speculative Sampling: Canonical Decomposition and Theoretical Limits
von: Khisti, Ashish, et al.
Veröffentlicht: (2024)
von: Khisti, Ashish, et al.
Veröffentlicht: (2024)
QA-Calibration of Language Model Confidence Scores
von: Manggala, Putra, et al.
Veröffentlicht: (2024)
von: Manggala, Putra, et al.
Veröffentlicht: (2024)
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
von: Maekawa, Aru, et al.
Veröffentlicht: (2024)
von: Maekawa, Aru, et al.
Veröffentlicht: (2024)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques
von: Kabir, Muhammad Rafsan, et al.
Veröffentlicht: (2024)
von: Kabir, Muhammad Rafsan, et al.
Veröffentlicht: (2024)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
von: Remy, François, et al.
Veröffentlicht: (2024)
von: Remy, François, et al.
Veröffentlicht: (2024)
The Geometry of Tokens in Internal Representations of Large Language Models
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
von: Fang, Luyang, et al.
Veröffentlicht: (2025)
von: Fang, Luyang, et al.
Veröffentlicht: (2025)
A Survey of On-Policy Distillation for Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Unsupervised Pretraining for Fact Verification by Language Model Distillation
von: Bazaga, Adrián, et al.
Veröffentlicht: (2023)
von: Bazaga, Adrián, et al.
Veröffentlicht: (2023)
Adversarial Moment-Matching Distillation of Large Language Models
von: Jia, Chen
Veröffentlicht: (2024)
von: Jia, Chen
Veröffentlicht: (2024)
On the Reasoning Abilities of Masked Diffusion Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2025)
von: Svete, Anej, et al.
Veröffentlicht: (2025)
Temporal Tokenization Strategies for Event Sequence Modeling with Large Language Models
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
Towards Linguistically-Aware and Language-Independent Tokenization for Large Language Models (LLMs)
von: Rahman, Abrar, et al.
Veröffentlicht: (2024)
von: Rahman, Abrar, et al.
Veröffentlicht: (2024)
Investigating Automatic Scoring and Feedback using Large Language Models
von: Katuka, Gloria Ashiya, et al.
Veröffentlicht: (2024)
von: Katuka, Gloria Ashiya, et al.
Veröffentlicht: (2024)
Efficient Temporal Tokenization for Mobility Prediction with Large Language Models
von: He, Haoyu, et al.
Veröffentlicht: (2025)
von: He, Haoyu, et al.
Veröffentlicht: (2025)
Rethinking Token Prediction: Tree-Structured Diffusion Language Model
von: Wu, Zihao, et al.
Veröffentlicht: (2026)
von: Wu, Zihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024) -
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
von: Phan, Buu, et al.
Veröffentlicht: (2024) -
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling
von: Phan, Buu, et al.
Veröffentlicht: (2025) -
List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
von: Rowan, Joseph, et al.
Veröffentlicht: (2025) -
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)