REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
Fuente:
arXiv
Guardado en:
| Autores principales: | Lasby, Mike, Lazarevich, Ivan, Sinnadurai, Nish, Lie, Sean, Ioannou, Yani, Thangarasa, Vithursan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SD$^2$: Self-Distilled Sparse Drafters
por: Lasby, Mike, et al.
Publicado: (2025)
por: Lasby, Mike, et al.
Publicado: (2025)
Self-Data Distillation for Recovering Quality in Pruned Large Language Models
por: Thangarasa, Vithursan, et al.
Publicado: (2024)
por: Thangarasa, Vithursan, et al.
Publicado: (2024)
TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
por: Sridhar, Aditya, et al.
Publicado: (2025)
por: Sridhar, Aditya, et al.
Publicado: (2025)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
por: Feng, Ruitao, et al.
Publicado: (2025)
por: Feng, Ruitao, et al.
Publicado: (2025)
Advancing Expert Specialization for Better MoE
por: Guo, Hongcan, et al.
Publicado: (2025)
por: Guo, Hongcan, et al.
Publicado: (2025)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
por: Vankov, Daniil, et al.
Publicado: (2026)
por: Vankov, Daniil, et al.
Publicado: (2026)
MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
por: Ganesan, Mugilan, et al.
Publicado: (2025)
por: Ganesan, Mugilan, et al.
Publicado: (2025)
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF
por: Song, Xingchen, et al.
Publicado: (2024)
por: Song, Xingchen, et al.
Publicado: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
GMoE: Empowering LLMs Fine-Tuning via MoE Graph Collaboration
por: Bai, Ting, et al.
Publicado: (2024)
por: Bai, Ting, et al.
Publicado: (2024)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
por: Taskooh, Hamidreza Kazemi, et al.
Publicado: (2026)
por: Taskooh, Hamidreza Kazemi, et al.
Publicado: (2026)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
por: Aslam, Nazia, et al.
Publicado: (2026)
por: Aslam, Nazia, et al.
Publicado: (2026)
SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning
por: Zou, Run, et al.
Publicado: (2026)
por: Zou, Run, et al.
Publicado: (2026)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
por: Wang, Xiaohua, et al.
Publicado: (2026)
por: Wang, Xiaohua, et al.
Publicado: (2026)
Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts
por: Martin, Liu O., et al.
Publicado: (2026)
por: Martin, Liu O., et al.
Publicado: (2026)
Active Few-Shot Learning for Text Classification
por: Ahmadnia, Saeed, et al.
Publicado: (2025)
por: Ahmadnia, Saeed, et al.
Publicado: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
por: Stacey, Joe, et al.
Publicado: (2026)
por: Stacey, Joe, et al.
Publicado: (2026)
Task Contamination: Language Models May Not Be Few-Shot Anymore
por: Li, Changmao, et al.
Publicado: (2023)
por: Li, Changmao, et al.
Publicado: (2023)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
por: Tian, Changxin, et al.
Publicado: (2025)
por: Tian, Changxin, et al.
Publicado: (2025)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
por: Yuan, Xin, et al.
Publicado: (2025)
por: Yuan, Xin, et al.
Publicado: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
por: Ewais, Ahmed, et al.
Publicado: (2026)
por: Ewais, Ahmed, et al.
Publicado: (2026)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
por: Akavarapu, V. S. D. S. Mahesh, et al.
Publicado: (2025)
por: Akavarapu, V. S. D. S. Mahesh, et al.
Publicado: (2025)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
por: Dinh, Tu Anh, et al.
Publicado: (2024)
por: Dinh, Tu Anh, et al.
Publicado: (2024)
Reducing Information Overload: Because Even Security Experts Need to Blink
por: Kuehn, Philipp, et al.
Publicado: (2022)
por: Kuehn, Philipp, et al.
Publicado: (2022)
Evaluating Large Language Models for Zero-Shot Disease Labeling in CT Radiology Reports Across Organ Systems
por: Garcia-Alcoser, Michael E., et al.
Publicado: (2025)
por: Garcia-Alcoser, Michael E., et al.
Publicado: (2025)
Super Apriel: One Checkpoint, Many Speeds
por: Labs, SLAM, et al.
Publicado: (2026)
por: Labs, SLAM, et al.
Publicado: (2026)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
por: Souza, Débora, et al.
Publicado: (2026)
por: Souza, Débora, et al.
Publicado: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
por: Schesch, Benedikt, et al.
Publicado: (2026)
por: Schesch, Benedikt, et al.
Publicado: (2026)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
por: Lv, Bo, et al.
Publicado: (2026)
por: Lv, Bo, et al.
Publicado: (2026)
Neural Machine Translation for Malayalam Paraphrase Generation
por: Varghese, Christeena, et al.
Publicado: (2024)
por: Varghese, Christeena, et al.
Publicado: (2024)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
por: Mamidanna, Siddarth, et al.
Publicado: (2025)
por: Mamidanna, Siddarth, et al.
Publicado: (2025)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
por: Park, Seungcheol, et al.
Publicado: (2023)
por: Park, Seungcheol, et al.
Publicado: (2023)
Ejemplares similares
-
SD$^2$: Self-Distilled Sparse Drafters
por: Lasby, Mike, et al.
Publicado: (2025) -
Self-Data Distillation for Recovering Quality in Pruned Large Language Models
por: Thangarasa, Vithursan, et al.
Publicado: (2024) -
TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
por: Sridhar, Aditya, et al.
Publicado: (2025) -
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
por: Feng, Ruitao, et al.
Publicado: (2025) -
Advancing Expert Specialization for Better MoE
por: Guo, Hongcan, et al.
Publicado: (2025)