LaDiMo: Layer-wise Distillation Inspired MoEfier
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Sungyoon, Kim, Youngjun, Moon, Kihyo, Jang, Minsung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ToDi: Token-wise Distillation via Fine-Grained Divergence Control
di: Jung, Seongryong, et al.
Pubblicazione: (2025)
di: Jung, Seongryong, et al.
Pubblicazione: (2025)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
di: Li, Zichong, et al.
Pubblicazione: (2025)
di: Li, Zichong, et al.
Pubblicazione: (2025)
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
di: Gurioli, Andrea, et al.
Pubblicazione: (2025)
di: Gurioli, Andrea, et al.
Pubblicazione: (2025)
On the Effect of Uncertainty on Layer-wise Inference Dynamics
di: Kim, Sunwoo, et al.
Pubblicazione: (2025)
di: Kim, Sunwoo, et al.
Pubblicazione: (2025)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
di: Son, Youngjun, et al.
Pubblicazione: (2025)
di: Son, Youngjun, et al.
Pubblicazione: (2025)
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
di: Byun, Hoyoon, et al.
Pubblicazione: (2025)
di: Byun, Hoyoon, et al.
Pubblicazione: (2025)
Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
di: Kim, Hazel, et al.
Pubblicazione: (2025)
di: Kim, Hazel, et al.
Pubblicazione: (2025)
Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
MoFE: Mixture of Frozen Experts Architecture
di: Seo, Jean, et al.
Pubblicazione: (2025)
di: Seo, Jean, et al.
Pubblicazione: (2025)
Fine-grained Gender Control in Machine Translation with Large Language Models
di: Lee, Minwoo, et al.
Pubblicazione: (2024)
di: Lee, Minwoo, et al.
Pubblicazione: (2024)
Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset
di: Yu, Hyojeong, et al.
Pubblicazione: (2026)
di: Yu, Hyojeong, et al.
Pubblicazione: (2026)
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
di: Kim, Sungnyun, et al.
Pubblicazione: (2025)
di: Kim, Sungnyun, et al.
Pubblicazione: (2025)
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
di: Kim, Mirae, et al.
Pubblicazione: (2026)
di: Kim, Mirae, et al.
Pubblicazione: (2026)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
di: Lee, Yunseung, et al.
Pubblicazione: (2026)
di: Lee, Yunseung, et al.
Pubblicazione: (2026)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
di: Bae, Sangmin, et al.
Pubblicazione: (2024)
di: Bae, Sangmin, et al.
Pubblicazione: (2024)
STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
di: Lee, Kyumin, et al.
Pubblicazione: (2025)
di: Lee, Kyumin, et al.
Pubblicazione: (2025)
MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning
di: Yang, Kang, et al.
Pubblicazione: (2025)
di: Yang, Kang, et al.
Pubblicazione: (2025)
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
di: Yang, Runyan, et al.
Pubblicazione: (2025)
di: Yang, Runyan, et al.
Pubblicazione: (2025)
Layer-wise Swapping for Generalizable Multilingual Safety
di: Shin, Hyunseo, et al.
Pubblicazione: (2026)
di: Shin, Hyunseo, et al.
Pubblicazione: (2026)
Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking
di: Kim, Joeun, et al.
Pubblicazione: (2026)
di: Kim, Joeun, et al.
Pubblicazione: (2026)
SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
di: Kim, Jungwoo, et al.
Pubblicazione: (2025)
di: Kim, Jungwoo, et al.
Pubblicazione: (2025)
State-of-the-Art Arabic Language Modeling with Sparse MoE Fine-Tuning and Chain-of-Thought Distillation
di: Singh, Navan Preet, et al.
Pubblicazione: (2026)
di: Singh, Navan Preet, et al.
Pubblicazione: (2026)
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
di: Wu, Haoyuan, et al.
Pubblicazione: (2025)
di: Wu, Haoyuan, et al.
Pubblicazione: (2025)
MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
di: Goyal, Agam, et al.
Pubblicazione: (2025)
di: Goyal, Agam, et al.
Pubblicazione: (2025)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
di: Kim, Minsung, et al.
Pubblicazione: (2025)
di: Kim, Minsung, et al.
Pubblicazione: (2025)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2025)
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2025)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
di: Tang, Yehui, et al.
Pubblicazione: (2025)
di: Tang, Yehui, et al.
Pubblicazione: (2025)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
di: Cao, Mingyu, et al.
Pubblicazione: (2024)
di: Cao, Mingyu, et al.
Pubblicazione: (2024)
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
di: Maekawa, Aru, et al.
Pubblicazione: (2024)
di: Maekawa, Aru, et al.
Pubblicazione: (2024)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
ExpGuard: LLM Content Moderation in Specialized Domains
di: Choi, Minseok, et al.
Pubblicazione: (2026)
di: Choi, Minseok, et al.
Pubblicazione: (2026)
LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring
di: Jang, Jinhee, et al.
Pubblicazione: (2025)
di: Jang, Jinhee, et al.
Pubblicazione: (2025)
Preference Distillation via Value based Reinforcement Learning
di: Kwon, Minchan, et al.
Pubblicazione: (2025)
di: Kwon, Minchan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ToDi: Token-wise Distillation via Fine-Grained Divergence Control
di: Jung, Seongryong, et al.
Pubblicazione: (2025) -
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
di: Yang, Mingyu, et al.
Pubblicazione: (2025) -
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
di: Li, Zichong, et al.
Pubblicazione: (2025) -
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
di: Gurioli, Andrea, et al.
Pubblicazione: (2025) -
On the Effect of Uncertainty on Layer-wise Inference Dynamics
di: Kim, Sunwoo, et al.
Pubblicazione: (2025)