Direct Neural Machine Translation with Task-level Mixture of Experts models
Fuente:
arXiv
Saved in:
| Main Authors: | Tourni, Isidora Chara, Naskar, Subhajit |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SandboxAQ's submission to MRL 2024 Shared Task on Multi-lingual Multi-task Information Retrieval
by: Tourni, Isidora Chara, et al.
Published: (2024)
by: Tourni, Isidora Chara, et al.
Published: (2024)
MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine Translation
by: Huang, Langlin, et al.
Published: (2024)
by: Huang, Langlin, et al.
Published: (2024)
Detecting Frames in News Headlines and Lead Images in U.S. Gun Violence Coverage
by: Tourni, Isidora Chara, et al.
Published: (2024)
by: Tourni, Isidora Chara, et al.
Published: (2024)
Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation
by: Ghassabi, Mehrdad, et al.
Published: (2026)
by: Ghassabi, Mehrdad, et al.
Published: (2026)
From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation
by: Barman, Amit, et al.
Published: (2025)
by: Barman, Amit, et al.
Published: (2025)
MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding Methods
by: Finkelstein, Mara, et al.
Published: (2023)
by: Finkelstein, Mara, et al.
Published: (2023)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Direct Preference Optimization for Neural Machine Translation with Minimum Bayes Risk Decoding
by: Yang, Guangyu, et al.
Published: (2023)
by: Yang, Guangyu, et al.
Published: (2023)
Task-Routed Mixture-of-Experts with Cognitive Appraisal for Implicit Sentiment Analysis
by: Chai, Yaping, et al.
Published: (2026)
by: Chai, Yaping, et al.
Published: (2026)
Machine Translation Models are Zero-Shot Detectors of Translation Direction
by: Wastl, Michelle, et al.
Published: (2024)
by: Wastl, Michelle, et al.
Published: (2024)
A Case Study on Context-Aware Neural Machine Translation with Multi-Task Learning
by: Appicharla, Ramakrishna, et al.
Published: (2024)
by: Appicharla, Ramakrishna, et al.
Published: (2024)
Direct Speech-to-Speech Neural Machine Translation: A Survey
by: Gupta, Mahendra, et al.
Published: (2024)
by: Gupta, Mahendra, et al.
Published: (2024)
Mixture of Neuron Experts
by: Cheng, Runxi, et al.
Published: (2025)
by: Cheng, Runxi, et al.
Published: (2025)
THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation
by: Liang, Yunlong, et al.
Published: (2025)
by: Liang, Yunlong, et al.
Published: (2025)
Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
by: Su, Guinan, et al.
Published: (2025)
by: Su, Guinan, et al.
Published: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
by: Deng, Guanzhi, et al.
Published: (2026)
by: Deng, Guanzhi, et al.
Published: (2026)
Evaluating Structural Generalization in Neural Machine Translation
by: Kumon, Ryoma, et al.
Published: (2024)
by: Kumon, Ryoma, et al.
Published: (2024)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
by: Ran, Junfeng, et al.
Published: (2025)
by: Ran, Junfeng, et al.
Published: (2025)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
Mixture of Lookup Experts
by: Jie, Shibo, et al.
Published: (2025)
by: Jie, Shibo, et al.
Published: (2025)
Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization
by: Uhlig, Kaden, et al.
Published: (2024)
by: Uhlig, Kaden, et al.
Published: (2024)
DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection
by: Ma, Guoxin, et al.
Published: (2025)
by: Ma, Guoxin, et al.
Published: (2025)
Mixture of Experts for Low-Resource LLMs
by: Joseph, Ori Bar, et al.
Published: (2026)
by: Joseph, Ori Bar, et al.
Published: (2026)
Layerwise Recurrent Router for Mixture-of-Experts
by: Qiu, Zihan, et al.
Published: (2024)
by: Qiu, Zihan, et al.
Published: (2024)
SAMoRA: Semantic-Aware Mixture of LoRA Experts for Task-Adaptive Learning
by: Shi, Boyan, et al.
Published: (2026)
by: Shi, Boyan, et al.
Published: (2026)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
by: Chen, Szu-Chi, et al.
Published: (2026)
by: Chen, Szu-Chi, et al.
Published: (2026)
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
by: Tan, Zheyue, et al.
Published: (2025)
by: Tan, Zheyue, et al.
Published: (2025)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
by: Wang, Yaoxiang, et al.
Published: (2025)
by: Wang, Yaoxiang, et al.
Published: (2025)
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
by: Guo, Hongcheng, et al.
Published: (2025)
by: Guo, Hongcheng, et al.
Published: (2025)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
by: Bai, Jun, et al.
Published: (2025)
by: Bai, Jun, et al.
Published: (2025)
Heterogeneous Encoders Scaling In The Transformer For Neural Machine Translation
by: Hu, Jia Cheng, et al.
Published: (2023)
by: Hu, Jia Cheng, et al.
Published: (2023)
Syntax-Aware Complex-Valued Neural Machine Translation
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
Top-down string-to-dependency Neural Machine Translation
by: Kondo, Shuhei, et al.
Published: (2026)
by: Kondo, Shuhei, et al.
Published: (2026)
Deterministic Reversible Data Augmentation for Neural Machine Translation
by: Yao, Jiashu, et al.
Published: (2024)
by: Yao, Jiashu, et al.
Published: (2024)
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
Evolving Knowledge Distillation for Lightweight Neural Machine Translation
by: Zhang, Xuewen, et al.
Published: (2026)
by: Zhang, Xuewen, et al.
Published: (2026)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Similar Items
-
SandboxAQ's submission to MRL 2024 Shared Task on Multi-lingual Multi-task Information Retrieval
by: Tourni, Isidora Chara, et al.
Published: (2024) -
MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine Translation
by: Huang, Langlin, et al.
Published: (2024) -
Detecting Frames in News Headlines and Lead Images in U.S. Gun Violence Coverage
by: Tourni, Isidora Chara, et al.
Published: (2024) -
Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation
by: Ghassabi, Mehrdad, et al.
Published: (2026) -
From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation
by: Barman, Amit, et al.
Published: (2025)