Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bai, Jun, Tong, Minghao, Liu, Yang, Jia, Zixia, Zheng, Zilong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915612555476992
author Bai, Jun
Tong, Minghao
Liu, Yang
Jia, Zixia
Zheng, Zilong
author_facet Bai, Jun
Tong, Minghao
Liu, Yang
Jia, Zixia
Zheng, Zilong
contents Context faithfulness is essential for reliable reasoning in context-dependent scenarios. However, large language models often struggle to ground their outputs in the provided context, resulting in irrelevant responses. Inspired by the emergent expert specialization observed in mixture-of-experts architectures, this work investigates whether certain experts exhibit specialization in context utilization, offering a potential pathway toward targeted optimization for improved context faithfulness. To explore this, we propose Router Lens, a method that accurately identifies context-faithful experts. Our analysis reveals that these experts progressively amplify attention to relevant contextual information, thereby enhancing context grounding. Building on this insight, we introduce Context-faithful Expert Fine-Tuning (CEFT), a lightweight optimization approach that selectively fine-tunes context-faithful experts. Experiments across a wide range of benchmarks and models demonstrate that CEFT matches or surpasses the performance of full fine-tuning while being significantly more efficient.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
Bai, Jun
Tong, Minghao
Liu, Yang
Jia, Zixia
Zheng, Zilong
Computation and Language
Context faithfulness is essential for reliable reasoning in context-dependent scenarios. However, large language models often struggle to ground their outputs in the provided context, resulting in irrelevant responses. Inspired by the emergent expert specialization observed in mixture-of-experts architectures, this work investigates whether certain experts exhibit specialization in context utilization, offering a potential pathway toward targeted optimization for improved context faithfulness. To explore this, we propose Router Lens, a method that accurately identifies context-faithful experts. Our analysis reveals that these experts progressively amplify attention to relevant contextual information, thereby enhancing context grounding. Building on this insight, we introduce Context-faithful Expert Fine-Tuning (CEFT), a lightweight optimization approach that selectively fine-tunes context-faithful experts. Experiments across a wide range of benchmarks and models demonstrate that CEFT matches or surpasses the performance of full fine-tuning while being significantly more efficient.
title Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
topic Computation and Language
url https://arxiv.org/abs/2508.19594