Investigating the potential of Sparse Mixtures-of-Experts for multi-domain neural machine translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chirkova, Nadezhda, Nikoulina, Vassilina, Meunier, Jean-Luc, Bérard, Alexandre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Key ingredients for effective zero-shot cross-lingual knowledge transfer in generative tasks
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
Zero-shot cross-lingual transfer in instruction tuning of large language models
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2023)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2023)
Retrieval-augmented generation in multilingual settings
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
Adapting Large Language Models for Multi-Domain Retrieval-Augmented-Generation
von: Misrahi, Alexandre, et al.
Veröffentlicht: (2025)
von: Misrahi, Alexandre, et al.
Veröffentlicht: (2025)
DiffLoRA: Differential Low-Rank Adapters for Large Language Models
von: Misrahi, Alexandre, et al.
Veröffentlicht: (2025)
von: Misrahi, Alexandre, et al.
Veröffentlicht: (2025)
Provence: efficient and robust context pruning for retrieval-augmented generation
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2025)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2025)
Retrieval-Augmented LLM Agents: Learning to Learn from Experience
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2026)
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2026)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
von: Rau, David, et al.
Veröffentlicht: (2024)
von: Rau, David, et al.
Veröffentlicht: (2024)
FrenchToxicityPrompts: a Large Benchmark for Evaluating and Mitigating Toxicity in French Texts
von: Brun, Caroline, et al.
Veröffentlicht: (2024)
von: Brun, Caroline, et al.
Veröffentlicht: (2024)
Toward domain-specific machine translation and quality estimation systems
von: Sharami, Javad Pourmostafa Roshan
Veröffentlicht: (2026)
von: Sharami, Javad Pourmostafa Roshan
Veröffentlicht: (2026)
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
von: Pham, Loc, et al.
Veröffentlicht: (2026)
von: Pham, Loc, et al.
Veröffentlicht: (2026)
Training Sparse Mixture Of Experts Text Embedding Models
von: Nussbaum, Zach, et al.
Veröffentlicht: (2025)
von: Nussbaum, Zach, et al.
Veröffentlicht: (2025)
LLM-as-a-qualitative-judge: automating error analysis in natural language generation
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2025)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2025)
SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
Can professional translators identify machine-generated text?
von: Farrell, Michael
Veröffentlicht: (2026)
von: Farrell, Michael
Veröffentlicht: (2026)
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
von: Lyu, Boxuan, et al.
Veröffentlicht: (2026)
von: Lyu, Boxuan, et al.
Veröffentlicht: (2026)
HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts
von: Zhong, Tao, et al.
Veröffentlicht: (2026)
von: Zhong, Tao, et al.
Veröffentlicht: (2026)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
Can postgraduate translation students identify machine-generated text?
von: Farrell, Michael
Veröffentlicht: (2025)
von: Farrell, Michael
Veröffentlicht: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
Unlocking Reasoning Capability on Machine Translation in Large Language Models
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
Automated evaluation of LLMs for effective machine translation of Mandarin Chinese to English
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Routing-Free Mixture-of-Experts
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
Multilingual Routing in Mixture-of-Experts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
Superposition in Transformers: A Novel Way of Building Mixture of Experts
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
von: Chaliah, Ayoub Ben, et al.
Veröffentlicht: (2024)
ConstitutionalExperts: Training a Mixture of Principle-based Prompts
von: Petridis, Savvas, et al.
Veröffentlicht: (2024)
von: Petridis, Savvas, et al.
Veröffentlicht: (2024)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
von: Liu, Boan, et al.
Veröffentlicht: (2023)
von: Liu, Boan, et al.
Veröffentlicht: (2023)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Key ingredients for effective zero-shot cross-lingual knowledge transfer in generative tasks
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024) -
Zero-shot cross-lingual transfer in instruction tuning of large language models
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024) -
Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2023) -
Retrieval-augmented generation in multilingual settings
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024) -
Adapting Large Language Models for Multi-Domain Retrieval-Augmented-Generation
von: Misrahi, Alexandre, et al.
Veröffentlicht: (2025)