Secret mixtures of experts inside your LLM
Fuente:
arXiv
Saved in:
| Main Author: | Boix-Adsera, Enric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024)
by: Boix-Adsera, Enric
Published: (2024)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
When can transformers reason with abstract symbols?
by: Boix-Adsera, Enric, et al.
Published: (2023)
by: Boix-Adsera, Enric, et al.
Published: (2023)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
by: Abbe, Emmanuel, et al.
Published: (2022)
by: Abbe, Emmanuel, et al.
Published: (2022)
Scaling physics-informed hard constraints with mixture-of-experts
by: Chalapathi, Nithin, et al.
Published: (2024)
by: Chalapathi, Nithin, et al.
Published: (2024)
EWMoE: An effective model for global weather forecasting with mixture-of-experts
by: Gan, Lihao, et al.
Published: (2024)
by: Gan, Lihao, et al.
Published: (2024)
Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
by: Luo, Beier, et al.
Published: (2025)
by: Luo, Beier, et al.
Published: (2025)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Data-Prep-Kit: getting your data ready for LLM application development
by: Wood, David, et al.
Published: (2024)
by: Wood, David, et al.
Published: (2024)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
Deep Ensembles Secretly Perform Empirical Bayes
by: Loaiza-Ganem, Gabriel, et al.
Published: (2025)
by: Loaiza-Ganem, Gabriel, et al.
Published: (2025)
AI Epidemiology: achieving explainable AI through expert oversight patterns
by: Tempest-Walters, Kit
Published: (2025)
by: Tempest-Walters, Kit
Published: (2025)
Flora: Low-Rank Adapters Are Secretly Gradient Compressors
by: Hao, Yongchang, et al.
Published: (2024)
by: Hao, Yongchang, et al.
Published: (2024)
Graders should cheat: privileged information enables expert-level automated evaluations
by: Zhou, Jin Peng, et al.
Published: (2025)
by: Zhou, Jin Peng, et al.
Published: (2025)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
by: Park, Jongseok, et al.
Published: (2026)
by: Park, Jongseok, et al.
Published: (2026)
Representation Convergence: Mutual Distillation is Secretly a Form of Regularization
by: Xie, Zhengpeng, et al.
Published: (2025)
by: Xie, Zhengpeng, et al.
Published: (2025)
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025)
by: Hu, Hengyuan, et al.
Published: (2025)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models
by: Soltaniani, Farnaz, et al.
Published: (2026)
by: Soltaniani, Farnaz, et al.
Published: (2026)
Capturing waste collection planning expert knowledge in a fitness function through preference learning
by: Díaz, Laura Fernández, et al.
Published: (2024)
by: Díaz, Laura Fernández, et al.
Published: (2024)
Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)
by: Miyashita, Hisashi
Published: (2026)
by: Miyashita, Hisashi
Published: (2026)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
by: Bhatia, Abhinav, et al.
Published: (2023)
by: Bhatia, Abhinav, et al.
Published: (2023)
Gradient-free variational learning with conditional mixture networks
by: Heins, Conor, et al.
Published: (2024)
by: Heins, Conor, et al.
Published: (2024)
Your Transformer is Secretly Linear
by: Razzhigaev, Anton, et al.
Published: (2024)
by: Razzhigaev, Anton, et al.
Published: (2024)
Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data
by: Wu, Weichang, et al.
Published: (2025)
by: Wu, Weichang, et al.
Published: (2025)
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
by: Singh, Prabhant, et al.
Published: (2025)
by: Singh, Prabhant, et al.
Published: (2025)
DEQuify your force field: More efficient simulations using deep equilibrium models
by: Burger, Andreas, et al.
Published: (2025)
by: Burger, Andreas, et al.
Published: (2025)
Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph
by: Liu, Ning, et al.
Published: (2026)
by: Liu, Ning, et al.
Published: (2026)
Stochastic stem bucking using mixture density neural networks
by: Schmiedel, Simon
Published: (2024)
by: Schmiedel, Simon
Published: (2024)
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
by: Wijk, Hjalmar, et al.
Published: (2024)
by: Wijk, Hjalmar, et al.
Published: (2024)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
by: Elhassan, Fay, et al.
Published: (2025)
by: Elhassan, Fay, et al.
Published: (2025)
Surrogate uncertainty estimation for your time series forecasting black-box: learn when to trust
by: Erlygin, Leonid, et al.
Published: (2023)
by: Erlygin, Leonid, et al.
Published: (2023)
Making AI Less "Thirsty": Uncovering and Addressing the Secret Water Footprint of AI Models
by: Li, Pengfei, et al.
Published: (2023)
by: Li, Pengfei, et al.
Published: (2023)
Knowledge Discovery using Unsupervised Cognition
by: Ibias, Alfredo, et al.
Published: (2024)
by: Ibias, Alfredo, et al.
Published: (2024)
AMiD: Knowledge Distillation for LLMs with $α$-mixture Assistant Distribution
by: Shin, Donghyeok, et al.
Published: (2025)
by: Shin, Donghyeok, et al.
Published: (2025)
Similar Items
-
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
by: Boix-Adsera, Enric, et al.
Published: (2025) -
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025) -
Towards a theory of model distillation
by: Boix-Adsera, Enric
Published: (2024) -
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025) -
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)