Mixture of Experts with Soft Nearest Neighbor Loss: Resolving Expert Collapse via Representation Disentanglement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agarap, Abien Fred, Azcarraga, Arnulfo P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
k-Winners-Take-All Ensemble Neural Network
von: Agarap, Abien Fred, et al.
Veröffentlicht: (2024)
von: Agarap, Abien Fred, et al.
Veröffentlicht: (2024)
Deep Learning using Rectified Linear Units (ReLU)
von: Agarap, Abien Fred
Veröffentlicht: (2018)
von: Agarap, Abien Fred
Veröffentlicht: (2018)
Apparent Age Estimation: Challenges and Outcomes
von: Go, Justin Rainier, et al.
Veröffentlicht: (2026)
von: Go, Justin Rainier, et al.
Veröffentlicht: (2026)
Nearest Neighbor Representations of Neurons
von: Kilic, Kordag Mehmet, et al.
Veröffentlicht: (2024)
von: Kilic, Kordag Mehmet, et al.
Veröffentlicht: (2024)
Nearest Neighbor Representations of Neural Circuits
von: Kilic, Kordag Mehmet, et al.
Veröffentlicht: (2024)
von: Kilic, Kordag Mehmet, et al.
Veröffentlicht: (2024)
Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts
von: Wong, Man Yung
Veröffentlicht: (2026)
von: Wong, Man Yung
Veröffentlicht: (2026)
A Gated Residual Kolmogorov-Arnold Networks for Mixtures of Experts
von: Inzirillo, Hugo, et al.
Veröffentlicht: (2024)
von: Inzirillo, Hugo, et al.
Veröffentlicht: (2024)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
MoEUT: Mixture-of-Experts Universal Transformers
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
Approximation Rates and VC-Dimension Bounds for (P)ReLU MLP Mixture of Experts
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2024)
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2024)
Partial Soft-Matching Distance for Neural Representational Comparison with Partial Unit Correspondence
von: Kapoor, Chaitanya, et al.
Veröffentlicht: (2026)
von: Kapoor, Chaitanya, et al.
Veröffentlicht: (2026)
FUSE: Measure-Theoretic Compact Fuzzy Set Representation for Taxonomy Expansion
von: Xu, Fred, et al.
Veröffentlicht: (2025)
von: Xu, Fred, et al.
Veröffentlicht: (2025)
EOE: Evolutionary Optimization of Experts for Training Language Models
von: Chen, Yingshi
Veröffentlicht: (2025)
von: Chen, Yingshi
Veröffentlicht: (2025)
Collapse-Free Prototype Readout Layer for Transformer Encoders
von: Cirrincione, Giansalvo, et al.
Veröffentlicht: (2026)
von: Cirrincione, Giansalvo, et al.
Veröffentlicht: (2026)
Learning Symbolic Model-Agnostic Loss Functions via Meta-Learning
von: Raymond, Christian, et al.
Veröffentlicht: (2022)
von: Raymond, Christian, et al.
Veröffentlicht: (2022)
Soft Quality-Diversity Optimization
von: Hedayatian, Saeed, et al.
Veröffentlicht: (2025)
von: Hedayatian, Saeed, et al.
Veröffentlicht: (2025)
Structural Equation-VAE: Disentangled Latent Representations for Tabular Data
von: Zhang, Ruiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiyu, et al.
Veröffentlicht: (2025)
Learning Mixture-of-Experts for General-Purpose Black-Box Discrete Optimization
von: Liu, Shengcai, et al.
Veröffentlicht: (2024)
von: Liu, Shengcai, et al.
Veröffentlicht: (2024)
Effective Regularization Through Loss-Function Metalearning
von: Gonzalez, Santiago, et al.
Veröffentlicht: (2020)
von: Gonzalez, Santiago, et al.
Veröffentlicht: (2020)
Conditional Finite Mixtures of Poisson Distributions for Context-Dependent Neural Correlations
von: Sokoloski, Sacha, et al.
Veröffentlicht: (2019)
von: Sokoloski, Sacha, et al.
Veröffentlicht: (2019)
Towards Faster k-Nearest-Neighbor Machine Translation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2023)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2023)
Directly Learning Stock Trading Strategies Through Profit Guided Loss Functions
von: Kar, Devroop, et al.
Veröffentlicht: (2025)
von: Kar, Devroop, et al.
Veröffentlicht: (2025)
Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
von: Ishikawa, Satoki, et al.
Veröffentlicht: (2024)
von: Ishikawa, Satoki, et al.
Veröffentlicht: (2024)
On the Universal Representation Property of Spiking Neural Networks
von: Hundrieser, Shayan, et al.
Veröffentlicht: (2025)
von: Hundrieser, Shayan, et al.
Veröffentlicht: (2025)
Recurrent Distance Filtering for Graph Representation Learning
von: Ding, Yuhui, et al.
Veröffentlicht: (2023)
von: Ding, Yuhui, et al.
Veröffentlicht: (2023)
Deep Neural Regression Collapse
von: Rangamani, Akshay, et al.
Veröffentlicht: (2026)
von: Rangamani, Akshay, et al.
Veröffentlicht: (2026)
Stochastic Forward-Forward Learning through Representational Dimensionality Compression
von: Zhu, Zhichao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhichao, et al.
Veröffentlicht: (2025)
Efficient Estimation of Unique Components in Independent Component Analysis by Matrix Representation
von: Matsuda, Yoshitatsu, et al.
Veröffentlicht: (2024)
von: Matsuda, Yoshitatsu, et al.
Veröffentlicht: (2024)
Representation Learning in a Decomposed Encoder Design for Bio-inspired Hebbian Learning
von: Jaziri, Achref, et al.
Veröffentlicht: (2023)
von: Jaziri, Achref, et al.
Veröffentlicht: (2023)
Time to Spike? Understanding the Representational Power of Spiking Neural Networks in Discrete Time
von: Nguyen, Duc Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Duc Anh, et al.
Veröffentlicht: (2025)
CARLA: Self-supervised Contrastive Representation Learning for Time Series Anomaly Detection
von: Darban, Zahra Zamanzadeh, et al.
Veröffentlicht: (2023)
von: Darban, Zahra Zamanzadeh, et al.
Veröffentlicht: (2023)
ABG-NAS: Adaptive Bayesian Genetic Neural Architecture Search for Graph Representation Learning
von: Wang, Sixuan, et al.
Veröffentlicht: (2025)
von: Wang, Sixuan, et al.
Veröffentlicht: (2025)
DDCL: Deep Dual Competitive Learning: A Differentiable End-to-End Framework for Unsupervised Prototype-Based Representation Learning
von: Cirrincione, Giansalvo
Veröffentlicht: (2026)
von: Cirrincione, Giansalvo
Veröffentlicht: (2026)
Improving Pareto Set Learning for Expensive Multi-objective Optimization via Stein Variational Hypernetworks
von: Nguyen, Minh-Duc, et al.
Veröffentlicht: (2024)
von: Nguyen, Minh-Duc, et al.
Veröffentlicht: (2024)
Transferring Core Knowledge via Learngenes
von: Feng, Fu, et al.
Veröffentlicht: (2024)
von: Feng, Fu, et al.
Veröffentlicht: (2024)
Pareto-Optimal Anytime Algorithms via Bayesian Racing
von: Wurth, Jonathan, et al.
Veröffentlicht: (2026)
von: Wurth, Jonathan, et al.
Veröffentlicht: (2026)
DALex: Lexicase-like Selection via Diverse Aggregation
von: Ni, Andrew, et al.
Veröffentlicht: (2024)
von: Ni, Andrew, et al.
Veröffentlicht: (2024)
Exploring the Improvement of Evolutionary Computation via Large Language Models
von: Cai, Jinyu, et al.
Veröffentlicht: (2024)
von: Cai, Jinyu, et al.
Veröffentlicht: (2024)
Online Reliable Anomaly Detection via Neuromorphic Sensing and Communications
von: Shiraishi, Junya, et al.
Veröffentlicht: (2025)
von: Shiraishi, Junya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
k-Winners-Take-All Ensemble Neural Network
von: Agarap, Abien Fred, et al.
Veröffentlicht: (2024) -
Deep Learning using Rectified Linear Units (ReLU)
von: Agarap, Abien Fred
Veröffentlicht: (2018) -
Apparent Age Estimation: Challenges and Outcomes
von: Go, Justin Rainier, et al.
Veröffentlicht: (2026) -
Nearest Neighbor Representations of Neurons
von: Kilic, Kordag Mehmet, et al.
Veröffentlicht: (2024) -
Nearest Neighbor Representations of Neural Circuits
von: Kilic, Kordag Mehmet, et al.
Veröffentlicht: (2024)