Q-Jamba: Quaternion-Native Hybrid State-Space Language Models with 3.4× Parameter Compression
Fuente:
Zenodo
Guardado en:
| Autor principal: | |
|---|---|
| Formato: | Recurso digital |
| Lenguaje: | inglés |
| Publicado: |
Zenodo
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866902301344530432 |
|---|---|
| author | Aleksej Dzigirej |
| author_facet | Aleksej Dzigirej |
| contents | <p>We introduce Q-Jamba, a family of quaternion-native language model architectures that achieve 3.4× parameter compression through Hamilton weight sharing while matching or exceeding standard transformer quality. All linear projections are replaced with QuaternionLinear layers that construct full m×n weight matrices from mn/4 learned parameters via the Hamilton block structure. We extend this principle to selective state spaces, proposing Q-Mamba — to our knowledge, the first SSM with Hamilton recurrence — and Q-Jamba, a hybrid that interleaves Q-Mamba blocks with quaternion attention.</p> <p>On a 9-task reasoning benchmark (n=5 seeds each), Q-Linear (547K params) significantly outperforms a parameter-matched standard transformer (559K params) with Cohen's d≈5.0. Q-Jamba 4:2 (813K params) achieves the lowest validation loss of all arms (0.421±0.004 vs. 0.506±0.016, p<10⁻⁴). On WikiText-2, Q-Linear matches a 3.4× larger standard model (BPC 2.102 vs. 2.105). A controlled dual-axis ablation reveals that structured coupling — not the algebraic rules of the quaternion algebra — drives these gains for feed-forward weights, while Hamilton algebra remains essential for recurrent state transitions. These findings emerge from 45 experiments on a single consumer GPU.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18673701 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Q-Jamba: Quaternion-Native Hybrid State-Space Language Models with 3.4× Parameter Compression Aleksej Dzigirej quaternion neural networks state-space models parameter compression language models LLM structured weight sharing Mamba hybrid architectures Hamilton product parameter efficiency <p>We introduce Q-Jamba, a family of quaternion-native language model architectures that achieve 3.4× parameter compression through Hamilton weight sharing while matching or exceeding standard transformer quality. All linear projections are replaced with QuaternionLinear layers that construct full m×n weight matrices from mn/4 learned parameters via the Hamilton block structure. We extend this principle to selective state spaces, proposing Q-Mamba — to our knowledge, the first SSM with Hamilton recurrence — and Q-Jamba, a hybrid that interleaves Q-Mamba blocks with quaternion attention.</p> <p>On a 9-task reasoning benchmark (n=5 seeds each), Q-Linear (547K params) significantly outperforms a parameter-matched standard transformer (559K params) with Cohen's d≈5.0. Q-Jamba 4:2 (813K params) achieves the lowest validation loss of all arms (0.421±0.004 vs. 0.506±0.016, p<10⁻⁴). On WikiText-2, Q-Linear matches a 3.4× larger standard model (BPC 2.102 vs. 2.105). A controlled dual-axis ablation reveals that structured coupling — not the algebraic rules of the quaternion algebra — drives these gains for feed-forward weights, while Hamilton algebra remains essential for recurrent state transitions. These findings emerge from 45 experiments on a single consumer GPU.</p> |
| title | Q-Jamba: Quaternion-Native Hybrid State-Space Language Models with 3.4× Parameter Compression |
| topic | quaternion neural networks state-space models parameter compression language models LLM structured weight sharing Mamba hybrid architectures Hamilton product parameter efficiency |
| url | https://doi.org/10.5281/zenodo.18673701 |