Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Liangwei Nathan, Zhang, Wei Emma, Guo, Mingyu, Maennel, Olaf, Chen, Weitong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Why Large Language Models Can Be Ineffective in Time Series Analysis: The Impact of Modality Alignment
by: Zheng, Liangwei Nathan, et al.
Published: (2024)
by: Zheng, Liangwei Nathan, et al.
Published: (2024)
Lifting Manifolds to Mitigate Pseudo-Alignment in LLM4TS
by: Zheng, Liangwei Nathan, et al.
Published: (2025)
by: Zheng, Liangwei Nathan, et al.
Published: (2025)
Probing Routing-Conditional Calibration in Attention-Residual Transformers
by: Liang, Wenhao, et al.
Published: (2026)
by: Liang, Wenhao, et al.
Published: (2026)
LLaDA-MoE: A Sparse MoE Diffusion Language Model
by: Zhu, Fengqi, et al.
Published: (2025)
by: Zhu, Fengqi, et al.
Published: (2025)
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
by: Yun, Sukwon, et al.
Published: (2024)
by: Yun, Sukwon, et al.
Published: (2024)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
by: Xu, Yu, et al.
Published: (2026)
by: Xu, Yu, et al.
Published: (2026)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
by: Farhat, Yehya, et al.
Published: (2023)
by: Farhat, Yehya, et al.
Published: (2023)
Kolmogorov-Arnold Networks (KAN) for Time Series Classification and Robust Analysis
by: Dong, Chang, et al.
Published: (2024)
by: Dong, Chang, et al.
Published: (2024)
Irregularity-Informed Time Series Analysis: Adaptive Modelling of Spatial and Temporal Dynamics
by: Zheng, Liangwei Nathan, et al.
Published: (2024)
by: Zheng, Liangwei Nathan, et al.
Published: (2024)
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
by: Liang, Wenhao, et al.
Published: (2025)
by: Liang, Wenhao, et al.
Published: (2025)
Free-Knots Kolmogorov-Arnold Network: On the Analysis of Spline Knots and Advancing Stability
by: Zheng, Liangwewi Nathan, et al.
Published: (2025)
by: Zheng, Liangwewi Nathan, et al.
Published: (2025)
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
by: Qin, Guangshuo, et al.
Published: (2026)
by: Qin, Guangshuo, et al.
Published: (2026)
GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts
by: Liao, Minwen, et al.
Published: (2025)
by: Liao, Minwen, et al.
Published: (2025)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
by: Huang, Zongle, et al.
Published: (2025)
by: Huang, Zongle, et al.
Published: (2025)
Devil in the Tail: A Multi-Modal Framework for Drug-Drug Interaction Prediction in Long Tail Distinction
by: Zheng, Liangwei Nathan, et al.
Published: (2024)
by: Zheng, Liangwei Nathan, et al.
Published: (2024)
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
by: Ma, Yingjie, et al.
Published: (2024)
by: Ma, Yingjie, et al.
Published: (2024)
CAMEL: Confidence-Gated Reflection for Reward Modeling
by: Zhu, Zirui, et al.
Published: (2026)
by: Zhu, Zirui, et al.
Published: (2026)
FFT-MoE: Efficient Federated Fine-Tuning for Foundation Models via Large-scale Sparse MoE under Heterogeneous Edge
by: Hu, Gang, et al.
Published: (2025)
by: Hu, Gang, et al.
Published: (2025)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
by: Hannah, Lauren. A, et al.
Published: (2025)
by: Hannah, Lauren. A, et al.
Published: (2025)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
by: Zhao, Jiayu, et al.
Published: (2026)
by: Zhao, Jiayu, et al.
Published: (2026)
PostHoc FREE Calibrating on Kolmogorov Arnold Networks
by: Liang, Wenhao, et al.
Published: (2025)
by: Liang, Wenhao, et al.
Published: (2025)
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
by: Zhao, Xinyuan, et al.
Published: (2026)
by: Zhao, Xinyuan, et al.
Published: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
by: Wang, Haodong, et al.
Published: (2025)
by: Wang, Haodong, et al.
Published: (2025)
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
by: Wu, Haoze, et al.
Published: (2024)
by: Wu, Haoze, et al.
Published: (2024)
Sparse-by-Design Cross-Modality Prediction: L0-Gated Representations for Reliable and Efficient Learning
by: Cenacchi, Filippo
Published: (2026)
by: Cenacchi, Filippo
Published: (2026)
Random Initialization of Gated Sparse Adapters
by: Retault, Vi, et al.
Published: (2025)
by: Retault, Vi, et al.
Published: (2025)
Collaborative Compression for Large-Scale MoE Deployment on Edge
by: Chen, Yixiao, et al.
Published: (2025)
by: Chen, Yixiao, et al.
Published: (2025)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
by: Chen, Yanlong, et al.
Published: (2026)
by: Chen, Yanlong, et al.
Published: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
by: Guo, Wentao, et al.
Published: (2025)
by: Guo, Wentao, et al.
Published: (2025)
DynamicGate MLP Conditional Computation via Learned Structural Dropout and Input Dependent Gating for Functional Plasticity
by: Choi, Yong Il
Published: (2026)
by: Choi, Yong Il
Published: (2026)
Sigma-MoE-Tiny Technical Report
by: Hu, Qingguo, et al.
Published: (2025)
by: Hu, Qingguo, et al.
Published: (2025)
DeepOmni: Towards Seamless and Smart Speech Interaction with Adaptive Modality-Specific MoE
by: Shao, Hang, et al.
Published: (2025)
by: Shao, Hang, et al.
Published: (2025)
The Confidence Gate Theorem: When Should Ranked Decision Systems Abstain?
by: Doku, Ronald
Published: (2026)
by: Doku, Ronald
Published: (2026)
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
by: Xu, Zukang, et al.
Published: (2026)
by: Xu, Zukang, et al.
Published: (2026)
Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
by: Guan, Xiaoliu, et al.
Published: (2025)
by: Guan, Xiaoliu, et al.
Published: (2025)
Confidence-Gated Robot Autonomy: When Does Uncertainty Actually Help?
by: Gaus, Johannes A., et al.
Published: (2026)
by: Gaus, Johannes A., et al.
Published: (2026)
E3x: $\mathrm{E}(3)$-Equivariant Deep Learning Made Easy
by: Unke, Oliver T., et al.
Published: (2024)
by: Unke, Oliver T., et al.
Published: (2024)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
by: Zhong, Zhengjia, et al.
Published: (2026)
by: Zhong, Zhengjia, et al.
Published: (2026)
Similar Items
-
Understanding Why Large Language Models Can Be Ineffective in Time Series Analysis: The Impact of Modality Alignment
by: Zheng, Liangwei Nathan, et al.
Published: (2024) -
Lifting Manifolds to Mitigate Pseudo-Alignment in LLM4TS
by: Zheng, Liangwei Nathan, et al.
Published: (2025) -
Probing Routing-Conditional Calibration in Attention-Residual Transformers
by: Liang, Wenhao, et al.
Published: (2026) -
LLaDA-MoE: A Sparse MoE Diffusion Language Model
by: Zhu, Fengqi, et al.
Published: (2025) -
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
by: Yun, Sukwon, et al.
Published: (2024)