Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Chuanyang, Sun, Jiankai, Gao, Yihang, Xie, Enze, Wang, Yuehao, Wang, Peihao, Xu, Ting, Chang, Matthew, Ren, Liliang, Li, Jingyao, Xiong, Jing, Rasul, Kashif, Schwager, Mac, Schneider, Anderson, Wang, Zhangyang, Nevmyvaka, Yuriy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cubit: Token Mixer with Kernel Ridge Regression
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
SAS: Simulated Attention Score
by: Zheng, Chuanyang, et al.
Published: (2025)
by: Zheng, Chuanyang, et al.
Published: (2025)
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
Improving Reasoning for Diffusion Language Models via Group Diffusion Policy Optimization
by: Rojas, Kevin, et al.
Published: (2025)
by: Rojas, Kevin, et al.
Published: (2025)
Structural Knowledge Informed Continual Multivariate Time Series Forecasting
by: Pan, Zijie, et al.
Published: (2024)
by: Pan, Zijie, et al.
Published: (2024)
Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
by: Roger, Alexis, et al.
Published: (2025)
by: Roger, Alexis, et al.
Published: (2025)
Speculative Sampling for Parametric Temporal Point Processes
by: Biloš, Marin, et al.
Published: (2025)
by: Biloš, Marin, et al.
Published: (2025)
Beyond Benign Overfitting in Nadaraya-Watson Interpolators
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
by: Hogan, Brendan R., et al.
Published: (2026)
by: Hogan, Brendan R., et al.
Published: (2026)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
by: Garg, Sahil, et al.
Published: (2024)
by: Garg, Sahil, et al.
Published: (2024)
Aria-NeRF: Multimodal Egocentric View Synthesis
by: Sun, Jiankai, et al.
Published: (2023)
by: Sun, Jiankai, et al.
Published: (2023)
Nadaraya-Watson kernel smoothing as a random energy model
by: Zavatone-Veth, Jacob A., et al.
Published: (2024)
by: Zavatone-Veth, Jacob A., et al.
Published: (2024)
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
by: Wang, Peihao, et al.
Published: (2025)
by: Wang, Peihao, et al.
Published: (2025)
A Mechanistic Study of Tabular Foundation Models
by: Biloš, Marin, et al.
Published: (2026)
by: Biloš, Marin, et al.
Published: (2026)
FAST-Splat: Fast, Ambiguity-Free Semantics Transfer in Gaussian Splatting
by: Shorinwa, Ola, et al.
Published: (2024)
by: Shorinwa, Ola, et al.
Published: (2024)
Nadaraya-Watson Type Estimator of the Transition Density Function for Diffusion Processes
by: Marie, Nicolas, et al.
Published: (2025)
by: Marie, Nicolas, et al.
Published: (2025)
TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
by: Ning, Kanghui, et al.
Published: (2025)
by: Ning, Kanghui, et al.
Published: (2025)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method
by: Zheng, Yan, et al.
Published: (2024)
by: Zheng, Yan, et al.
Published: (2024)
Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing
by: Wang, Peihao, et al.
Published: (2024)
by: Wang, Peihao, et al.
Published: (2024)
Position: Weight Space Should Be a First-Class Generative AI Modality
by: Wang, Zhangyang, et al.
Published: (2026)
by: Wang, Zhangyang, et al.
Published: (2026)
Technical Report: Full-Stack Fine-Tuning for the Q Programming Language
by: Hogan, Brendan R., et al.
Published: (2025)
by: Hogan, Brendan R., et al.
Published: (2025)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
by: Riachi, Roland, et al.
Published: (2025)
by: Riachi, Roland, et al.
Published: (2025)
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
by: Shorinwa, Ola, et al.
Published: (2025)
by: Shorinwa, Ola, et al.
Published: (2025)
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity
by: Fokoué, Ernest
Published: (2026)
by: Fokoué, Ernest
Published: (2026)
Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision
by: Ning, Kanghui, et al.
Published: (2025)
by: Ning, Kanghui, et al.
Published: (2025)
Privacy Amplification by Structured Subsampling for Deep Differentially Private Time Series Forecasting
by: Schuchardt, Jan, et al.
Published: (2025)
by: Schuchardt, Jan, et al.
Published: (2025)
$\textbf{S}^2$IP-LLM: Semantic Space Informed Prompt Learning with LLM for Time Series Forecasting
by: Pan, Zijie, et al.
Published: (2024)
by: Pan, Zijie, et al.
Published: (2024)
FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
by: Liu, Hengyu, et al.
Published: (2025)
by: Liu, Hengyu, et al.
Published: (2025)
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
by: Huang, Suning, et al.
Published: (2025)
by: Huang, Suning, et al.
Published: (2025)
Forecasting with Hyper-Trees
by: März, Alexander, et al.
Published: (2024)
by: März, Alexander, et al.
Published: (2024)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
by: Zhan, Zheng, et al.
Published: (2025)
by: Zhan, Zheng, et al.
Published: (2025)
DAPE V2: Process Attention Score as Feature Map for Length Extrapolation
by: Zheng, Chuanyang, et al.
Published: (2024)
by: Zheng, Chuanyang, et al.
Published: (2024)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
Reweighting Improves Conditional Risk Bounds
by: Zhang, Yikai, et al.
Published: (2025)
by: Zhang, Yikai, et al.
Published: (2025)
Empowering Time Series Analysis with Large Language Models: A Survey
by: Jiang, Yushan, et al.
Published: (2024)
by: Jiang, Yushan, et al.
Published: (2024)
Recurrent Interpolants for Probabilistic Time Series Prediction
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Similar Items
-
Cubit: Token Mixer with Kernel Ridge Regression
by: Zheng, Chuanyang, et al.
Published: (2026) -
GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization
by: Zheng, Chuanyang, et al.
Published: (2026) -
SAS: Simulated Attention Score
by: Zheng, Chuanyang, et al.
Published: (2025) -
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
by: Zhao, Jinze, et al.
Published: (2024) -
Improving Reasoning for Diffusion Language Models via Group Diffusion Policy Optimization
by: Rojas, Kevin, et al.
Published: (2025)