Rethinking LLM Ensembling from the Perspective of Mixture Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Jiale, Jiang, Yuchu, Wu, Peijun, Liu, Chonghan, Zhou, Joey Tianyi, Yang, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
d$^2$Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
von: Jiang, Yuchu, et al.
Veröffentlicht: (2025)
von: Jiang, Yuchu, et al.
Veröffentlicht: (2025)
Rethinking Data Mixing from the Perspective of Large Language Models
von: Xu, Yuanjian, et al.
Veröffentlicht: (2026)
von: Xu, Yuanjian, et al.
Veröffentlicht: (2026)
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?
von: Li, Wenzhe, et al.
Veröffentlicht: (2025)
von: Li, Wenzhe, et al.
Veröffentlicht: (2025)
Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
von: Liu, Yexiang, et al.
Veröffentlicht: (2025)
von: Liu, Yexiang, et al.
Veröffentlicht: (2025)
DiscoUQ: Structured Disagreement Analysis for Uncertainty Quantification in LLM Agent Ensembles
von: Jiang, Bo
Veröffentlicht: (2026)
von: Jiang, Bo
Veröffentlicht: (2026)
LLM-TOPLA: Efficient LLM Ensemble by Maximising Diversity
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
LLM Router: Rethinking Routing with Prefill Activations
von: Varshney, Tanay, et al.
Veröffentlicht: (2026)
von: Varshney, Tanay, et al.
Veröffentlicht: (2026)
Rethinking Local Learning: A Cheaper and Faster Recipe for LLM Post-Training
von: Shi, Hengyu, et al.
Veröffentlicht: (2026)
von: Shi, Hengyu, et al.
Veröffentlicht: (2026)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
Rethinking Machine Unlearning for Large Language Models
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
Rethinking Token Prediction: Tree-Structured Diffusion Language Model
von: Wu, Zihao, et al.
Veröffentlicht: (2026)
von: Wu, Zihao, et al.
Veröffentlicht: (2026)
SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture
von: Han, Jiayi, et al.
Veröffentlicht: (2024)
von: Han, Jiayi, et al.
Veröffentlicht: (2024)
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
von: Xu, Qingshu, et al.
Veröffentlicht: (2025)
von: Xu, Qingshu, et al.
Veröffentlicht: (2025)
Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective
von: Zhang, Yanan, et al.
Veröffentlicht: (2024)
von: Zhang, Yanan, et al.
Veröffentlicht: (2024)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
Rethinking Token Reduction for State Space Models
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
von: Huang, Yangyi, et al.
Veröffentlicht: (2026)
von: Huang, Yangyi, et al.
Veröffentlicht: (2026)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
von: Qiu, Wenjie, et al.
Veröffentlicht: (2025)
von: Qiu, Wenjie, et al.
Veröffentlicht: (2025)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
von: Zeng, Hanqing, et al.
Veröffentlicht: (2025)
von: Zeng, Hanqing, et al.
Veröffentlicht: (2025)
Flatter Tokens are More Valuable for Speculative Draft Model Training
von: Fan, Jiaming, et al.
Veröffentlicht: (2026)
von: Fan, Jiaming, et al.
Veröffentlicht: (2026)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
A Closer Look into Mixture-of-Experts in Large Language Models
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
RevOrder: A Novel Method for Enhanced Arithmetic in Language Models
von: Shen, Si, et al.
Veröffentlicht: (2024)
von: Shen, Si, et al.
Veröffentlicht: (2024)
CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
von: Feng, Jinyuan, et al.
Veröffentlicht: (2025)
von: Feng, Jinyuan, et al.
Veröffentlicht: (2025)
OMoE: Diversifying Mixture of Low-Rank Adaptation by Orthogonal Finetuning
von: Feng, Jinyuan, et al.
Veröffentlicht: (2025)
von: Feng, Jinyuan, et al.
Veröffentlicht: (2025)
Rethinking LLM Memorization through the Lens of Adversarial Compression
von: Schwarzschild, Avi, et al.
Veröffentlicht: (2024)
von: Schwarzschild, Avi, et al.
Veröffentlicht: (2024)
A Survey on Mixture of Experts in Large Language Models
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
GeoReasoner: Reasoning On Geospatially Grounded Context For Natural Language Understanding
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
von: Wang, An, et al.
Veröffentlicht: (2024)
von: Wang, An, et al.
Veröffentlicht: (2024)
Critical Data Size of Language Models from a Grokking Perspective
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
Kinetics: Rethinking Test-Time Scaling Laws
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
von: Jiang, Yuxian, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025) -
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
von: Li, Ziyue, et al.
Veröffentlicht: (2024) -
d$^2$Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
von: Jiang, Yuchu, et al.
Veröffentlicht: (2025) -
Rethinking Data Mixing from the Perspective of Large Language Models
von: Xu, Yuanjian, et al.
Veröffentlicht: (2026) -
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?
von: Li, Wenzhe, et al.
Veröffentlicht: (2025)