Saved in:
Bibliographic Details
Main Authors: Bae, Sangmin, Acun, Bilge, Lin, Chien-Yu, Habeeb, Haroun, Kim, Seungyeon, Luo, Liang, Wang, Junjie, Wu, Carole-Jean
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.04800
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911609688948736
author Bae, Sangmin
Acun, Bilge
Lin, Chien-Yu
Habeeb, Haroun
Kim, Seungyeon
Luo, Liang
Wang, Junjie
Wu, Carole-Jean
author_facet Bae, Sangmin
Acun, Bilge
Lin, Chien-Yu
Habeeb, Haroun
Kim, Seungyeon
Luo, Liang
Wang, Junjie
Wu, Carole-Jean
contents Recent progress in large language models demonstrates that hybrid architectures--combining self-attention mechanisms with structured state space models like Mamba--can achieve a compelling balance between modeling quality and computational efficiency, particularly for long-context tasks. While these hybrid models show promising performance, systematic comparisons of hybridization strategies and analyses on the key factors behind their effectiveness have not been clearly shared to the community. In this work, we present a holistic evaluation of hybrid architectures based on inter-layer (sequential) or intra-layer (parallel) fusion. We comprehensively evaluate these designs across multiple dimensions: language modeling and downstream task performance, long-context capabilities, scaling analysis, and training and inference efficiency. By investigating the core characteristics of their computational primitive, we identify the most critical elements for each hybridization strategy and further propose optimal design recipes for hybrid models. Our comprehensive analysis provides practical guidance and valuable insights for developing hybrid language models, facilitating the optimization of architectural configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04800
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
Bae, Sangmin
Acun, Bilge
Lin, Chien-Yu
Habeeb, Haroun
Kim, Seungyeon
Luo, Liang
Wang, Junjie
Wu, Carole-Jean
Computation and Language
Recent progress in large language models demonstrates that hybrid architectures--combining self-attention mechanisms with structured state space models like Mamba--can achieve a compelling balance between modeling quality and computational efficiency, particularly for long-context tasks. While these hybrid models show promising performance, systematic comparisons of hybridization strategies and analyses on the key factors behind their effectiveness have not been clearly shared to the community. In this work, we present a holistic evaluation of hybrid architectures based on inter-layer (sequential) or intra-layer (parallel) fusion. We comprehensively evaluate these designs across multiple dimensions: language modeling and downstream task performance, long-context capabilities, scaling analysis, and training and inference efficiency. By investigating the core characteristics of their computational primitive, we identify the most critical elements for each hybridization strategy and further propose optimal design recipes for hybrid models. Our comprehensive analysis provides practical guidance and valuable insights for developing hybrid language models, facilitating the optimization of architectural configurations.
title Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
topic Computation and Language
url https://arxiv.org/abs/2510.04800