Comba: Improving Bilinear RNNs with Closed-loop Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Jiaxi, Pan, Yongqi, Du, Jusen, Lan, Disen, Tang, Xiaqiang, Wen, Qingsong, Liang, Yuxuan, Sun, Weigao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912745003155456
author Hu, Jiaxi
Pan, Yongqi
Du, Jusen
Lan, Disen
Tang, Xiaqiang
Wen, Qingsong
Liang, Yuxuan
Sun, Weigao
author_facet Hu, Jiaxi
Pan, Yongqi
Du, Jusen
Lan, Disen
Tang, Xiaqiang
Wen, Qingsong
Liang, Yuxuan
Sun, Weigao
contents Recent efficient sequence modeling methods such as Gated DeltaNet, TTT, and RWKV-7 have achieved performance improvements by supervising the recurrent memory management through Delta learning rule. Unlike previous state-space models (e.g., Mamba) and gated linear attentions (e.g., GLA), these models introduce interactions between the recurrent state and the key vector, structurally resembling bilinear systems. In this paper, we first introduce the concept of Bilinear RNNs with a comprehensive analysis on the advantages and limitations of these models. Then, based on closed-loop control theory, we propose a novel Bilinear RNN variant named Comba, which adopts a scalar-plus-low-rank state transition, with both state feedback and output feedback corrections. We also implement a hardware-efficient chunk-wise parallel kernel in Triton and train models with 340M/1.3B parameters on large-scale corpus. Comba demonstrates superior performance and computation efficiency in both language and vision modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comba: Improving Bilinear RNNs with Closed-loop Control
Hu, Jiaxi
Pan, Yongqi
Du, Jusen
Lan, Disen
Tang, Xiaqiang
Wen, Qingsong
Liang, Yuxuan
Sun, Weigao
Machine Learning
Computation and Language
Recent efficient sequence modeling methods such as Gated DeltaNet, TTT, and RWKV-7 have achieved performance improvements by supervising the recurrent memory management through Delta learning rule. Unlike previous state-space models (e.g., Mamba) and gated linear attentions (e.g., GLA), these models introduce interactions between the recurrent state and the key vector, structurally resembling bilinear systems. In this paper, we first introduce the concept of Bilinear RNNs with a comprehensive analysis on the advantages and limitations of these models. Then, based on closed-loop control theory, we propose a novel Bilinear RNN variant named Comba, which adopts a scalar-plus-low-rank state transition, with both state feedback and output feedback corrections. We also implement a hardware-efficient chunk-wise parallel kernel in Triton and train models with 340M/1.3B parameters on large-scale corpus. Comba demonstrates superior performance and computation efficiency in both language and vision modeling.
title Comba: Improving Bilinear RNNs with Closed-loop Control
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2506.02475