Discussion: Effective and Interpretable Outcome Prediction by Training Sparse Mixtures of Linear Experts
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Folino, Francesco, Pontieri, Luigi, Sabatino, Pietro |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
par: Pham, Quang, et autres
Publié: (2024)
par: Pham, Quang, et autres
Publié: (2024)
Generating the Traces You Need: A Conditional Generative Model for Process Mining Data
par: Graziosi, Riccardo, et autres
Publié: (2024)
par: Graziosi, Riccardo, et autres
Publié: (2024)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
par: Chowdhury, Mohammed Nowaz Rabbani, et autres
Publié: (2024)
par: Chowdhury, Mohammed Nowaz Rabbani, et autres
Publié: (2024)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
par: Nikolic, Strahinja, et autres
Publié: (2025)
par: Nikolic, Strahinja, et autres
Publié: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
par: Panda, Ashwinee, et autres
Publié: (2025)
par: Panda, Ashwinee, et autres
Publié: (2025)
Combining Abstract Argumentation and Machine Learning for Efficiently Analyzing Low-Level Process Event Streams
par: Fazzinga, Bettina, et autres
Publié: (2025)
par: Fazzinga, Bettina, et autres
Publié: (2025)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
par: Nguyen, Dung V., et autres
Publié: (2025)
par: Nguyen, Dung V., et autres
Publié: (2025)
Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
par: Pavlitska, Svetlana, et autres
Publié: (2025)
par: Pavlitska, Svetlana, et autres
Publié: (2025)
Mixture of Concept Bottleneck Experts
par: De Santis, Francesco, et autres
Publié: (2026)
par: De Santis, Francesco, et autres
Publié: (2026)
Interpretable Cascading Mixture-of-Experts for Urban Traffic Congestion Prediction
par: Jiang, Wenzhao, et autres
Publié: (2024)
par: Jiang, Wenzhao, et autres
Publié: (2024)
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
par: Pan, Xinglin, et autres
Publié: (2025)
par: Pan, Xinglin, et autres
Publié: (2025)
On the Role of Discrete Representation in Sparse Mixture of Experts
par: Do, Giang, et autres
Publié: (2024)
par: Do, Giang, et autres
Publié: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
par: Pan, Bowen, et autres
Publié: (2024)
par: Pan, Bowen, et autres
Publié: (2024)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
par: Nakamura, Taishi, et autres
Publié: (2025)
par: Nakamura, Taishi, et autres
Publié: (2025)
Mixture of Experts Made Intrinsically Interpretable
par: Yang, Xingyi, et autres
Publié: (2025)
par: Yang, Xingyi, et autres
Publié: (2025)
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
par: Nguyen, Duc Anh, et autres
Publié: (2025)
par: Nguyen, Duc Anh, et autres
Publié: (2025)
Improving Routing in Sparse Mixture of Experts with Graph of Tokens
par: Nguyen, Tam, et autres
Publié: (2025)
par: Nguyen, Tam, et autres
Publié: (2025)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
par: Tran, Viet-Hoang, et autres
Publié: (2025)
par: Tran, Viet-Hoang, et autres
Publié: (2025)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
par: Chen, Shengzhuang, et autres
Publié: (2025)
par: Chen, Shengzhuang, et autres
Publié: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
par: Zhang, Zeliang, et autres
Publié: (2024)
par: Zhang, Zeliang, et autres
Publié: (2024)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
par: Ahrac, Sagi, et autres
Publié: (2026)
par: Ahrac, Sagi, et autres
Publié: (2026)
$ϕ$-Balancing for Mixture-of-Experts Training
par: Chen, Lizhang, et autres
Publié: (2026)
par: Chen, Lizhang, et autres
Publié: (2026)
CoSMoEs: Compact Sparse Mixture of Experts
par: Huber, Patrick, et autres
Publié: (2025)
par: Huber, Patrick, et autres
Publié: (2025)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
par: Cai, Weilin, et autres
Publié: (2024)
par: Cai, Weilin, et autres
Publié: (2024)
Combining Euclidean and Hyperbolic Representations for Node-level Anomaly Detection
par: Mungari, Simone, et autres
Publié: (2025)
par: Mungari, Simone, et autres
Publié: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
par: Muzio, Alexandre, et autres
Publié: (2024)
par: Muzio, Alexandre, et autres
Publié: (2024)
Prediction-powered Inference by Mixture of Experts
par: Gu, Yanwu, et autres
Publié: (2026)
par: Gu, Yanwu, et autres
Publié: (2026)
From Sparse to Soft Mixtures of Experts
par: Puigcerver, Joan, et autres
Publié: (2023)
par: Puigcerver, Joan, et autres
Publié: (2023)
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
par: Zhao, Jinze, et autres
Publié: (2024)
par: Zhao, Jinze, et autres
Publié: (2024)
Split-on-Share: Mixture of Sparse Experts for Task-Agnostic Continual Learning
par: Siddika, Fatema, et autres
Publié: (2026)
par: Siddika, Fatema, et autres
Publié: (2026)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
par: Nguyen, Huy, et autres
Publié: (2023)
par: Nguyen, Huy, et autres
Publié: (2023)
MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction
par: Aparício, Carolina, et autres
Publié: (2025)
par: Aparício, Carolina, et autres
Publié: (2025)
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
par: Morrison, Jacob, et autres
Publié: (2026)
par: Morrison, Jacob, et autres
Publié: (2026)
Self-Augmented Mixture-of-Experts for QoS Prediction
par: Cai, Kecheng, et autres
Publié: (2026)
par: Cai, Kecheng, et autres
Publié: (2026)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
par: Rastegar, Reza
Publié: (2026)
par: Rastegar, Reza
Publié: (2026)
Algebraformer: A Neural Approach to Linear Systems
par: Sittoni, Pietro, et autres
Publié: (2025)
par: Sittoni, Pietro, et autres
Publié: (2025)
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
par: Liu, Enshu, et autres
Publié: (2024)
par: Liu, Enshu, et autres
Publié: (2024)
Super-Linear: A Lightweight Pretrained Mixture of Linear Experts for Time Series Forecasting
par: Nochumsohn, Liran, et autres
Publié: (2025)
par: Nochumsohn, Liran, et autres
Publié: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
par: Jiang, Yukun, et autres
Publié: (2026)
par: Jiang, Yukun, et autres
Publié: (2026)
Documents similaires
-
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
par: Pham, Quang, et autres
Publié: (2024) -
Generating the Traces You Need: A Conditional Generative Model for Process Mining Data
par: Graziosi, Riccardo, et autres
Publié: (2024) -
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
par: Chowdhury, Mohammed Nowaz Rabbani, et autres
Publié: (2024) -
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
par: Nikolic, Strahinja, et autres
Publié: (2025) -
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
par: Panda, Ashwinee, et autres
Publié: (2025)