Hydra: A Modular Architecture for Efficient Long-Context Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaudhary, Siddharth, Patel, Dev, Chaudhary, Maheep, Browning, Bennett
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908598359031808
author Chaudhary, Siddharth
Patel, Dev
Chaudhary, Maheep
Browning, Bennett
author_facet Chaudhary, Siddharth
Patel, Dev
Chaudhary, Maheep
Browning, Bennett
contents The quadratic complexity of transformers fundamentally limits reasoning system deployment in resource-constrained and long-context settings. We introduce Hydra, a modular architecture based upon a state-space backbone which adaptively routes between complementary efficiency mechanisms: sparse global attention, mixture-of-experts, and dual memories comprising a reasoning workspace and product key memory. We evaluate a 29M parameter model measuring logical chaining accuracy and throughput on synthetic sequences, plus throughput on WikiText. Ablation studies use component-specific synthetic datasets to isolate individual mechanisms. Hydra achieves $3.01\times$ and $3.0\times$ throughput gains at 8K tokens for synthetic and WikiText datasets, respectively, and $10\times$ accuracy improvements on multi-step logical composition compared to equal-sized transformers. Ablations confirm each component's contribution: sparse attention captures long-range dependencies, experts specialize to input domains, and product key memory enables selective retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15099
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hydra: A Modular Architecture for Efficient Long-Context Reasoning
Chaudhary, Siddharth
Patel, Dev
Chaudhary, Maheep
Browning, Bennett
Machine Learning
Artificial Intelligence
The quadratic complexity of transformers fundamentally limits reasoning system deployment in resource-constrained and long-context settings. We introduce Hydra, a modular architecture based upon a state-space backbone which adaptively routes between complementary efficiency mechanisms: sparse global attention, mixture-of-experts, and dual memories comprising a reasoning workspace and product key memory. We evaluate a 29M parameter model measuring logical chaining accuracy and throughput on synthetic sequences, plus throughput on WikiText. Ablation studies use component-specific synthetic datasets to isolate individual mechanisms. Hydra achieves $3.01\times$ and $3.0\times$ throughput gains at 8K tokens for synthetic and WikiText datasets, respectively, and $10\times$ accuracy improvements on multi-step logical composition compared to equal-sized transformers. Ablations confirm each component's contribution: sparse attention captures long-range dependencies, experts specialize to input domains, and product key memory enables selective retrieval.
title Hydra: A Modular Architecture for Efficient Long-Context Reasoning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.15099