Bayesian Mixture-of-Experts: Towards Making LLMs Know What They Don't Know
Fuente:
arXiv
Saved in:
| Main Author: | Li, Albus Yizhuo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers
by: Li, Albus Yizhuo, et al.
Published: (2026)
by: Li, Albus Yizhuo, et al.
Published: (2026)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
by: Johnson, Daniel D., et al.
Published: (2024)
by: Johnson, Daniel D., et al.
Published: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024)
by: Kapoor, Sanyam, et al.
Published: (2024)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
by: Hariri, Mohsen, et al.
Published: (2025)
by: Hariri, Mohsen, et al.
Published: (2025)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
by: Pal, Arka, et al.
Published: (2025)
by: Pal, Arka, et al.
Published: (2025)
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
by: Dong, Zihan, et al.
Published: (2026)
by: Dong, Zihan, et al.
Published: (2026)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
by: Yan, Fanqi, et al.
Published: (2026)
by: Yan, Fanqi, et al.
Published: (2026)
Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
by: He, Langzhou, et al.
Published: (2025)
by: He, Langzhou, et al.
Published: (2025)
Can Bayesian Neural Networks Make Confident Predictions?
by: Fisher, Katharine, et al.
Published: (2025)
by: Fisher, Katharine, et al.
Published: (2025)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
Bayesian Semiparametric Mixture Cure (Frailty) Models
by: Kızılaslan, Fatih, et al.
Published: (2025)
by: Kızılaslan, Fatih, et al.
Published: (2025)
Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures
by: Thai, Tuan, et al.
Published: (2025)
by: Thai, Tuan, et al.
Published: (2025)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency without Model Sweeps
by: Hai, Do Tien, et al.
Published: (2025)
by: Hai, Do Tien, et al.
Published: (2025)
Fast Model Selection and Stable Optimization for Softmax-Gated Multinomial-Logistic Mixture of Experts Models
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
by: Cohen, Roi, et al.
Published: (2024)
by: Cohen, Roi, et al.
Published: (2024)
Towards Bayesian Data Selection
by: Rodemann, Julian
Published: (2024)
by: Rodemann, Julian
Published: (2024)
Show Me What You Don't Know: Efficient Sampling from Invariant Sets for Model Validation
by: Rousselot, Armand, et al.
Published: (2026)
by: Rousselot, Armand, et al.
Published: (2026)
What Makes Treatment Effects Identifiable? Characterizations and Estimators Beyond Unconfoundedness
by: Cai, Yang, et al.
Published: (2025)
by: Cai, Yang, et al.
Published: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Local Minima Structures in Gaussian Mixture Models
by: Chen, Yudong, et al.
Published: (2020)
by: Chen, Yudong, et al.
Published: (2020)
First, Learn What You Don't Know: Active Information Gathering for Driving at the Limits of Handling
by: Davydov, Alexander, et al.
Published: (2024)
by: Davydov, Alexander, et al.
Published: (2024)
Mixture of multilayer stochastic block models for multiview clustering
by: De Santiago, Kylliann, et al.
Published: (2024)
by: De Santiago, Kylliann, et al.
Published: (2024)
Deep Bayesian Inversion
by: Adler, Jonas, et al.
Published: (2018)
by: Adler, Jonas, et al.
Published: (2018)
Nonparametric MLE for Gaussian Location Mixtures: Certified Computation and Generic Behavior
by: Polyanskiy, Yury, et al.
Published: (2025)
by: Polyanskiy, Yury, et al.
Published: (2025)
Sharp Inequalities between Total Variation and Hellinger Distances for Gaussian Mixtures
by: Jung, Joonhyuk, et al.
Published: (2026)
by: Jung, Joonhyuk, et al.
Published: (2026)
Variational Approach for Efficient KL Divergence Estimation in Dirichlet Mixture Models
by: Pal, Samyajoy, et al.
Published: (2024)
by: Pal, Samyajoy, et al.
Published: (2024)
Adaptive Mean Estimation in the Hidden Markov sub-Gaussian Mixture Model
by: Karagulyan, Vahe, et al.
Published: (2024)
by: Karagulyan, Vahe, et al.
Published: (2024)
Gaussian Mixture Model with unknown diagonal covariances via continuous sparse regularization
by: Giard, Romane, et al.
Published: (2025)
by: Giard, Romane, et al.
Published: (2025)
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Sequential Bayesian Neural Subnetwork Ensembles
by: Jantre, Sanket, et al.
Published: (2022)
by: Jantre, Sanket, et al.
Published: (2022)
On uncertainty-penalized Bayesian information criterion
by: Thanasutives, Pongpisit, et al.
Published: (2024)
by: Thanasutives, Pongpisit, et al.
Published: (2024)
Online Statistical Inference in Decision-Making with Matrix Context
by: Han, Qiyu, et al.
Published: (2022)
by: Han, Qiyu, et al.
Published: (2022)
Singular Fluctuation as Specific Heat in Bayesian Learning
by: Plummer, Sean
Published: (2025)
by: Plummer, Sean
Published: (2025)
Thermodynamic Response Functions in Singular Bayesian Models
by: Plummer, Sean
Published: (2026)
by: Plummer, Sean
Published: (2026)
Frequentist Guarantees of Distributed (Non)-Bayesian Inference
by: Wu, Bohan, et al.
Published: (2023)
by: Wu, Bohan, et al.
Published: (2023)
Imagining What We Don't Know
by: Samuels, Lisa
Published: (2026)
by: Samuels, Lisa
Published: (2026)
Imagining What We Don't Know
by: Samuels, Lisa
Published: (2026)
by: Samuels, Lisa
Published: (2026)
Similar Items
-
Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers
by: Li, Albus Yizhuo, et al.
Published: (2026) -
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
by: Johnson, Daniel D., et al.
Published: (2024) -
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024) -
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025) -
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)