KD4MT: A Survey of Knowledge Distillation for Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | de Gibert, Ona, Attieh, Joseph, Mickus, Timothee, Scherrer, Yves, Tiedemann, Jörg |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Life Cycle-Aware Evaluation of Knowledge Distillation for Machine Translation: Environmental Impact and Translation Quality Trade-offs
by: Attieh, Joseph, et al.
Published: (2026)
by: Attieh, Joseph, et al.
Published: (2026)
I Have an Attention Bridge to Sell You: Generalization Capabilities of Modular Translation Architectures
by: Mickus, Timothee, et al.
Published: (2024)
by: Mickus, Timothee, et al.
Published: (2024)
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki
by: Mickus, Timothee, et al.
Published: (2024)
by: Mickus, Timothee, et al.
Published: (2024)
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
by: de Gibert, Ona, et al.
Published: (2025)
by: de Gibert, Ona, et al.
Published: (2025)
Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?
by: Ji, Shaoxiong, et al.
Published: (2024)
by: Ji, Shaoxiong, et al.
Published: (2024)
Isotropy, Clusters, and Classifiers
by: Mickus, Timothee, et al.
Published: (2024)
by: Mickus, Timothee, et al.
Published: (2024)
Open Machine Translation for Esperanto
by: de Gibert, Ona, et al.
Published: (2026)
by: de Gibert, Ona, et al.
Published: (2026)
A Comparison of Language Modeling and Translation as Multilingual Pretraining Objectives
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
DocHPLT: A Massively Multilingual Document-Level Translation Dataset
by: O'Brien, Dayyán, et al.
Published: (2025)
by: O'Brien, Dayyán, et al.
Published: (2025)
Why do Large Language Models Fail in Low-resource Translation? Unraveling the Token Dynamics of Large Language Models for Machine Translation
by: Qian, Shenbin, et al.
Published: (2026)
by: Qian, Shenbin, et al.
Published: (2026)
SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes
by: Vázquez, Raúl, et al.
Published: (2025)
by: Vázquez, Raúl, et al.
Published: (2025)
The Impact of Vocabulary Overlaps on Knowledge Transfer in Multilingual Machine Translation
by: Itkonen, Oona, et al.
Published: (2026)
by: Itkonen, Oona, et al.
Published: (2026)
Adapting Definition Modeling for New Languages: A Case Study on Belarusian
by: Kazakouskaya, Daniela, et al.
Published: (2025)
by: Kazakouskaya, Daniela, et al.
Published: (2025)
Your Model is Overconfident, and Other Lies We Tell Ourselves
by: Mickus, Timothee, et al.
Published: (2025)
by: Mickus, Timothee, et al.
Published: (2025)
MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine Translation
by: Li, Jiahuan, et al.
Published: (2024)
by: Li, Jiahuan, et al.
Published: (2024)
SemEval-2024 Shared Task 6: SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes
by: Mickus, Timothee, et al.
Published: (2024)
by: Mickus, Timothee, et al.
Published: (2024)
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models
by: Luo, Hengyu, et al.
Published: (2025)
by: Luo, Hengyu, et al.
Published: (2025)
Domain-specific or Uncertainty-aware models: Does it really make a difference for biomedical text classification?
by: Sinha, Aman, et al.
Published: (2024)
by: Sinha, Aman, et al.
Published: (2024)
Evolving Knowledge Distillation for Lightweight Neural Machine Translation
by: Zhang, Xuewen, et al.
Published: (2026)
by: Zhang, Xuewen, et al.
Published: (2026)
Test-Time Scaling of Reasoning Models for Machine Translation
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
GOSt-MT: A Knowledge Graph for Occupation-related Gender Biases in Machine Translation
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2024)
Grounded Satirical Generation with RAG
by: Itkonen, Oona, et al.
Published: (2026)
by: Itkonen, Oona, et al.
Published: (2026)
Self-Evolution Knowledge Distillation for LLM-based Machine Translation
by: Song, Yuncheng, et al.
Published: (2024)
by: Song, Yuncheng, et al.
Published: (2024)
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation
by: Huang, Chenyang, et al.
Published: (2025)
by: Huang, Chenyang, et al.
Published: (2025)
$\mathcal{X}$-KD: General Experiential Knowledge Distillation for Large Language Models
by: Cai, Yuang, et al.
Published: (2026)
by: Cai, Yuang, et al.
Published: (2026)
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
by: Kadlčík, Marek, et al.
Published: (2025)
by: Kadlčík, Marek, et al.
Published: (2025)
AXOLOTL'24 Shared Task on Multilingual Explainable Semantic Change Modeling
by: Fedorova, Mariia, et al.
Published: (2024)
by: Fedorova, Mariia, et al.
Published: (2024)
BridG MT: Enhancing LLMs' Machine Translation Capabilities with Sentence Bridging and Gradual MT
by: Choi, Seung-Woo, et al.
Published: (2024)
by: Choi, Seung-Woo, et al.
Published: (2024)
Memorization Inheritance in Sequence-Level Knowledge Distillation for Neural Machine Translation
by: Dankers, Verna, et al.
Published: (2025)
by: Dankers, Verna, et al.
Published: (2025)
Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation
by: Zhang, Songming, et al.
Published: (2023)
by: Zhang, Songming, et al.
Published: (2023)
DSG-KD: Knowledge Distillation from Domain-Specific to General Language Models
by: Cho, Sangyeon, et al.
Published: (2024)
by: Cho, Sangyeon, et al.
Published: (2024)
Definition generation for lexical semantic change detection
by: Fedorova, Mariia, et al.
Published: (2024)
by: Fedorova, Mariia, et al.
Published: (2024)
Explainability of machine learning approaches in forensic linguistics: a case study in geolinguistic authorship profiling
by: Roemling, Dana, et al.
Published: (2024)
by: Roemling, Dana, et al.
Published: (2024)
DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation
by: Vu, Duc Trung, et al.
Published: (2026)
by: Vu, Duc Trung, et al.
Published: (2026)
ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation
by: Li, Kunquan, et al.
Published: (2026)
by: Li, Kunquan, et al.
Published: (2026)
GrammaMT: Improving Machine Translation with Grammar-Informed In-Context Learning
by: Ramos, Rita, et al.
Published: (2024)
by: Ramos, Rita, et al.
Published: (2024)
MT-LENS: An all-in-one Toolkit for Better Machine Translation Evaluation
by: Gilabert, Javier García, et al.
Published: (2024)
by: Gilabert, Javier García, et al.
Published: (2024)
CULL-MT: Compression Using Language and Layer pruning for Machine Translation
by: Rostami, Pedram, et al.
Published: (2024)
by: Rostami, Pedram, et al.
Published: (2024)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Omnilingual MT: Machine Translation for 1,600 Languages
by: Omnilingual MT Team, et al.
Published: (2026)
by: Omnilingual MT Team, et al.
Published: (2026)
Similar Items
-
Life Cycle-Aware Evaluation of Knowledge Distillation for Machine Translation: Environmental Impact and Translation Quality Trade-offs
by: Attieh, Joseph, et al.
Published: (2026) -
I Have an Attention Bridge to Sell You: Generalization Capabilities of Modular Translation Architectures
by: Mickus, Timothee, et al.
Published: (2024) -
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki
by: Mickus, Timothee, et al.
Published: (2024) -
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
by: de Gibert, Ona, et al.
Published: (2025) -
Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?
by: Ji, Shaoxiong, et al.
Published: (2024)