Zamba: A Compact 7B SSM Hybrid Model
Fuente:
arXiv
Saved in:
| Main Authors: | Glorioso, Paolo, Anthony, Quentin, Tokpanov, Yury, Whittington, James, Pilault, Jonathan, Ibrahim, Adam, Millidge, Beren |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Zamba2 Suite: Technical Report
by: Glorioso, Paolo, et al.
Published: (2024)
by: Glorioso, Paolo, et al.
Published: (2024)
Zyda: A 1.3T Dataset for Open Language Modeling
by: Tokpanov, Yury, et al.
Published: (2024)
by: Tokpanov, Yury, et al.
Published: (2024)
BlackMamba: Mixture of Experts for State-Space Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Zyda-2: a 5 Trillion Token High-Quality Dataset
by: Tokpanov, Yury, et al.
Published: (2024)
by: Tokpanov, Yury, et al.
Published: (2024)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
by: Shyam, Vasudev, et al.
Published: (2024)
by: Shyam, Vasudev, et al.
Published: (2024)
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
by: Figliolia, Tomas, et al.
Published: (2025)
by: Figliolia, Tomas, et al.
Published: (2025)
Hybrid Associative Memories
by: Lufkin, Leon, et al.
Published: (2026)
by: Lufkin, Leon, et al.
Published: (2026)
ZAYA1-8B Technical Report
by: Washbourne, Robert, et al.
Published: (2026)
by: Washbourne, Robert, et al.
Published: (2026)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
by: Alonso, Nick, et al.
Published: (2024)
by: Alonso, Nick, et al.
Published: (2024)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
by: Ibrahim, Adam, et al.
Published: (2024)
by: Ibrahim, Adam, et al.
Published: (2024)
ZUNA: Flexible EEG Superresolution with Position-Aware Diffusion Autoencoders
by: Warner, Christopher, et al.
Published: (2026)
by: Warner, Christopher, et al.
Published: (2026)
Generalising E-prop to Deep Networks
by: Millidge, Beren
Published: (2025)
by: Millidge, Beren
Published: (2025)
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
by: Anthony, Quentin, et al.
Published: (2025)
by: Anthony, Quentin, et al.
Published: (2025)
LongSSM: On the Length Extension of State-space Models in Language Modelling
by: Wang, Shida
Published: (2024)
by: Wang, Shida
Published: (2024)
Heracles: A Hybrid SSM-Transformer Model for High-Resolution Image and Time-Series Analysis
by: Patro, Badri N., et al.
Published: (2024)
by: Patro, Badri N., et al.
Published: (2024)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
by: Bick, Aviv, et al.
Published: (2026)
by: Bick, Aviv, et al.
Published: (2026)
ZAYA1-VL-8B Technical Report
by: Shapourian, Hassan, et al.
Published: (2026)
by: Shapourian, Hassan, et al.
Published: (2026)
Exploring Action-Centric Representations Through the Lens of Rate-Distortion Theory
by: Varona, Miguel de Llanza, et al.
Published: (2024)
by: Varona, Miguel de Llanza, et al.
Published: (2024)
Specialised or Generic? Tokenization Choices for Radiology Language Models
by: Warr, Hermione, et al.
Published: (2025)
by: Warr, Hermione, et al.
Published: (2025)
Distilling LLMs' Decomposition Abilities into Compact Language Models
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
Compact Language Models via Pruning and Knowledge Distillation
by: Muralidharan, Saurav, et al.
Published: (2024)
by: Muralidharan, Saurav, et al.
Published: (2024)
Causal Abstraction in Model Interpretability: A Compact Survey
by: Zhang, Yihao
Published: (2024)
by: Zhang, Yihao
Published: (2024)
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG
by: Alonso, Nicholas, et al.
Published: (2024)
by: Alonso, Nicholas, et al.
Published: (2024)
OTCE: Hybrid SSM and Attention with Cross Domain Mixture of Experts to construct Observer-Thinker-Conceiver-Expresser
by: Shi, Jingze, et al.
Published: (2024)
by: Shi, Jingze, et al.
Published: (2024)
EmbedLLM: Learning Compact Representations of Large Language Models
by: Zhuang, Richard, et al.
Published: (2024)
by: Zhuang, Richard, et al.
Published: (2024)
Simple Policy Gradients for Reasoning with Diffusion Language Models
by: Zhan, Anthony
Published: (2025)
by: Zhan, Anthony
Published: (2025)
ELM: A Hybrid Ensemble of Language Models for Automated Tumor Group Classification in Population-Based Cancer Registries
by: Gondara, Lovedeep, et al.
Published: (2025)
by: Gondara, Lovedeep, et al.
Published: (2025)
Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch
by: Pfister, Jan, et al.
Published: (2024)
by: Pfister, Jan, et al.
Published: (2024)
CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Language Models
by: Lakkapragada, Venkat Akhil
Published: (2026)
by: Lakkapragada, Venkat Akhil
Published: (2026)
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
by: Ali, Mehdi, et al.
Published: (2024)
by: Ali, Mehdi, et al.
Published: (2024)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
by: Microsoft, et al.
Published: (2025)
by: Microsoft, et al.
Published: (2025)
Internal narratives parameterise affective states
by: Onysk, Jakub, et al.
Published: (2025)
by: Onysk, Jakub, et al.
Published: (2025)
Trillion 7B Technical Report
by: Han, Sungjun, et al.
Published: (2025)
by: Han, Sungjun, et al.
Published: (2025)
KV Cache Transform Coding for Compact Storage in LLM Inference
by: Staniszewski, Konrad, et al.
Published: (2025)
by: Staniszewski, Konrad, et al.
Published: (2025)
A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding Networks
by: Salvatori, Tommaso, et al.
Published: (2022)
by: Salvatori, Tommaso, et al.
Published: (2022)
Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
Baichuan2-Sum: Instruction Finetune Baichuan2-7B Model for Dialogue Summarization
by: Xiao, Jianfei, et al.
Published: (2024)
by: Xiao, Jianfei, et al.
Published: (2024)
Similar Items
-
The Zamba2 Suite: Technical Report
by: Glorioso, Paolo, et al.
Published: (2024) -
Zyda: A 1.3T Dataset for Open Language Modeling
by: Tokpanov, Yury, et al.
Published: (2024) -
BlackMamba: Mixture of Experts for State-Space Models
by: Anthony, Quentin, et al.
Published: (2024) -
Zyda-2: a 5 Trillion Token High-Quality Dataset
by: Tokpanov, Yury, et al.
Published: (2024) -
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
by: Shyam, Vasudev, et al.
Published: (2024)