The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Herbst, Jeremy, Wermter, Stefan, Lee, Jae Hee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
von: Codefuse, et al.
Veröffentlicht: (2025)
von: Codefuse, et al.
Veröffentlicht: (2025)
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
von: Wang, Shouren, et al.
Veröffentlicht: (2026)
von: Wang, Shouren, et al.
Veröffentlicht: (2026)
Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts
von: Martin, Liu O., et al.
Veröffentlicht: (2026)
von: Martin, Liu O., et al.
Veröffentlicht: (2026)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
von: Avinash, Mynampati Sri Ranganadha
Veröffentlicht: (2026)
von: Avinash, Mynampati Sri Ranganadha
Veröffentlicht: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
von: Vankov, Daniil, et al.
Veröffentlicht: (2026)
von: Vankov, Daniil, et al.
Veröffentlicht: (2026)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
von: Gao, Yutong, et al.
Veröffentlicht: (2026)
von: Gao, Yutong, et al.
Veröffentlicht: (2026)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
von: Yordanov, Yordan, et al.
Veröffentlicht: (2026)
von: Yordanov, Yordan, et al.
Veröffentlicht: (2026)
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
von: Lasby, Mike, et al.
Veröffentlicht: (2025)
von: Lasby, Mike, et al.
Veröffentlicht: (2025)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
von: Koo, Ryan, et al.
Veröffentlicht: (2023)
von: Koo, Ryan, et al.
Veröffentlicht: (2023)
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
von: Lee, Myeonghwa, et al.
Veröffentlicht: (2024)
von: Lee, Myeonghwa, et al.
Veröffentlicht: (2024)
Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict
von: Venkata, Pruthvinath Jeripity
Veröffentlicht: (2026)
von: Venkata, Pruthvinath Jeripity
Veröffentlicht: (2026)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
von: Nainani, Jatin, et al.
Veröffentlicht: (2024)
von: Nainani, Jatin, et al.
Veröffentlicht: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
Integrating Expert Labels into LLM-based Emission Goal Detection: Example Selection vs Automatic Prompt Design
von: Wrzalik, Marco, et al.
Veröffentlicht: (2024)
von: Wrzalik, Marco, et al.
Veröffentlicht: (2024)
Random-Set Large Language Models
von: Mubashar, Muhammad, et al.
Veröffentlicht: (2025)
von: Mubashar, Muhammad, et al.
Veröffentlicht: (2025)
Improving Language Models with Intentional Analysis
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
Danoliteracy of Generative Large Language Models
von: Holm, Søren Vejlgaard, et al.
Veröffentlicht: (2024)
von: Holm, Søren Vejlgaard, et al.
Veröffentlicht: (2024)
SWI: Speaking with Intent in Large Language Models
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
Uncovering Biases with Reflective Large Language Models
von: Chang, Edward Y.
Veröffentlicht: (2024)
von: Chang, Edward Y.
Veröffentlicht: (2024)
Uncovering Latent Human Wellbeing in Language Model Embeddings
von: Freire, Pedro, et al.
Veröffentlicht: (2024)
von: Freire, Pedro, et al.
Veröffentlicht: (2024)
Self-Supervised Position Debiasing for Large Language Models
von: Liu, Zhongkun, et al.
Veröffentlicht: (2024)
von: Liu, Zhongkun, et al.
Veröffentlicht: (2024)
Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
von: Zhang, Liyi, et al.
Veröffentlicht: (2025)
von: Zhang, Liyi, et al.
Veröffentlicht: (2025)
HAMMER: Hamiltonian Curiosity Augmented Large Language Model Reinforcement
von: Yang, Ming, et al.
Veröffentlicht: (2025)
von: Yang, Ming, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
von: Saini, Mayank, et al.
Veröffentlicht: (2025)
von: Saini, Mayank, et al.
Veröffentlicht: (2025)
RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models
von: Lin, Tianqianjin, et al.
Veröffentlicht: (2025)
von: Lin, Tianqianjin, et al.
Veröffentlicht: (2025)
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
von: Quevedo, Ernesto, et al.
Veröffentlicht: (2024)
von: Quevedo, Ernesto, et al.
Veröffentlicht: (2024)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
von: Gilhuly, Colleen, et al.
Veröffentlicht: (2025)
von: Gilhuly, Colleen, et al.
Veröffentlicht: (2025)
External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
von: Yan, Jian
Veröffentlicht: (2025)
von: Yan, Jian
Veröffentlicht: (2025)
Concept Navigation and Classification via Open-Source Large Language Model Processing
von: Kubli, Maël
Veröffentlicht: (2025)
von: Kubli, Maël
Veröffentlicht: (2025)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
von: Song, Chenyang, et al.
Veröffentlicht: (2024)
von: Song, Chenyang, et al.
Veröffentlicht: (2024)
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
von: Chen, Haolin, et al.
Veröffentlicht: (2024)
von: Chen, Haolin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
von: Codefuse, et al.
Veröffentlicht: (2025) -
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
von: Wang, Shouren, et al.
Veröffentlicht: (2026) -
Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts
von: Martin, Liu O., et al.
Veröffentlicht: (2026) -
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
von: Avinash, Mynampati Sri Ranganadha
Veröffentlicht: (2026) -
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)