The Zamba2 Suite: Technical Report
Fuente:
arXiv
Guardado en:
| Autores principales: | Glorioso, Paolo, Anthony, Quentin, Tokpanov, Yury, Golubeva, Anna, Shyam, Vasudev, Whittington, James, Pilault, Jonathan, Millidge, Beren |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Zamba: A Compact 7B SSM Hybrid Model
por: Glorioso, Paolo, et al.
Publicado: (2024)
por: Glorioso, Paolo, et al.
Publicado: (2024)
Zyda: A 1.3T Dataset for Open Language Modeling
por: Tokpanov, Yury, et al.
Publicado: (2024)
por: Tokpanov, Yury, et al.
Publicado: (2024)
BlackMamba: Mixture of Experts for State-Space Models
por: Anthony, Quentin, et al.
Publicado: (2024)
por: Anthony, Quentin, et al.
Publicado: (2024)
Zyda-2: a 5 Trillion Token High-Quality Dataset
por: Tokpanov, Yury, et al.
Publicado: (2024)
por: Tokpanov, Yury, et al.
Publicado: (2024)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
por: Shyam, Vasudev, et al.
Publicado: (2024)
por: Shyam, Vasudev, et al.
Publicado: (2024)
ZAYA1-8B Technical Report
por: Washbourne, Robert, et al.
Publicado: (2026)
por: Washbourne, Robert, et al.
Publicado: (2026)
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
por: Figliolia, Tomas, et al.
Publicado: (2025)
por: Figliolia, Tomas, et al.
Publicado: (2025)
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
por: Anthony, Quentin, et al.
Publicado: (2025)
por: Anthony, Quentin, et al.
Publicado: (2025)
ZAYA1-VL-8B Technical Report
por: Shapourian, Hassan, et al.
Publicado: (2026)
por: Shapourian, Hassan, et al.
Publicado: (2026)
Hybrid Associative Memories
por: Lufkin, Leon, et al.
Publicado: (2026)
por: Lufkin, Leon, et al.
Publicado: (2026)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
por: Alonso, Nick, et al.
Publicado: (2024)
por: Alonso, Nick, et al.
Publicado: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
por: Shyam, Vasu, et al.
Publicado: (2026)
por: Shyam, Vasu, et al.
Publicado: (2026)
ZUNA: Flexible EEG Superresolution with Position-Aware Diffusion Autoencoders
por: Warner, Christopher, et al.
Publicado: (2026)
por: Warner, Christopher, et al.
Publicado: (2026)
Generalising E-prop to Deep Networks
por: Millidge, Beren
Publicado: (2025)
por: Millidge, Beren
Publicado: (2025)
PLaMo 2 Technical Report
por: Networks, Preferred, et al.
Publicado: (2025)
por: Networks, Preferred, et al.
Publicado: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
por: Guldimann, Philipp, et al.
Publicado: (2024)
por: Guldimann, Philipp, et al.
Publicado: (2024)
Nemotron-4 15B Technical Report
por: Parmar, Jupinder, et al.
Publicado: (2024)
por: Parmar, Jupinder, et al.
Publicado: (2024)
PCMind-2.1-Kaiyuan-2B Technical Report
por: Luo, Kairong, et al.
Publicado: (2025)
por: Luo, Kairong, et al.
Publicado: (2025)
Yi-Lightning Technical Report
por: Wake, Alan, et al.
Publicado: (2024)
por: Wake, Alan, et al.
Publicado: (2024)
Nemotron-4 340B Technical Report
por: Nvidia, et al.
Publicado: (2024)
por: Nvidia, et al.
Publicado: (2024)
Trillion 7B Technical Report
por: Han, Sungjun, et al.
Publicado: (2025)
por: Han, Sungjun, et al.
Publicado: (2025)
Skywork Open Reasoner 1 Technical Report
por: He, Jujie, et al.
Publicado: (2025)
por: He, Jujie, et al.
Publicado: (2025)
EuroLLM-22B: Technical Report
por: Ramos, Miguel Moura, et al.
Publicado: (2026)
por: Ramos, Miguel Moura, et al.
Publicado: (2026)
EuroLLM-9B: Technical Report
por: Martins, Pedro Henrique, et al.
Publicado: (2025)
por: Martins, Pedro Henrique, et al.
Publicado: (2025)
Technical Report: Small Language Model for Japanese Clinical and Medicine
por: Watanabe, Shogo
Publicado: (2024)
por: Watanabe, Shogo
Publicado: (2024)
Latxa: An Open Language Model and Evaluation Suite for Basque
por: Etxaniz, Julen, et al.
Publicado: (2024)
por: Etxaniz, Julen, et al.
Publicado: (2024)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
por: Yang, An, et al.
Publicado: (2024)
por: Yang, An, et al.
Publicado: (2024)
Ovis2.5 Technical Report
por: Lu, Shiyin, et al.
Publicado: (2025)
por: Lu, Shiyin, et al.
Publicado: (2025)
Kwai Summary Attention Technical Report
por: Chu, Chenglong, et al.
Publicado: (2026)
por: Chu, Chenglong, et al.
Publicado: (2026)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
por: Rohekar, Raanan Y., et al.
Publicado: (2024)
por: Rohekar, Raanan Y., et al.
Publicado: (2024)
Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
por: Dade, Nii Osae Osae, et al.
Publicado: (2025)
por: Dade, Nii Osae Osae, et al.
Publicado: (2025)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
Exploring Action-Centric Representations Through the Lens of Rate-Distortion Theory
por: Varona, Miguel de Llanza, et al.
Publicado: (2024)
por: Varona, Miguel de Llanza, et al.
Publicado: (2024)
TTS-1 Technical Report
por: Atamanenko, Oleg, et al.
Publicado: (2025)
por: Atamanenko, Oleg, et al.
Publicado: (2025)
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG
por: Alonso, Nicholas, et al.
Publicado: (2024)
por: Alonso, Nicholas, et al.
Publicado: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
por: Kim, Eunsu, et al.
Publicado: (2025)
por: Kim, Eunsu, et al.
Publicado: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
por: Gohil, Vasudev
Publicado: (2025)
por: Gohil, Vasudev
Publicado: (2025)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
por: Iyer, Vivek, et al.
Publicado: (2025)
por: Iyer, Vivek, et al.
Publicado: (2025)
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
por: Varshney, Prasoon, et al.
Publicado: (2025)
por: Varshney, Prasoon, et al.
Publicado: (2025)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
por: Hwang, Bokwang, et al.
Publicado: (2025)
por: Hwang, Bokwang, et al.
Publicado: (2025)
Ejemplares similares
-
Zamba: A Compact 7B SSM Hybrid Model
por: Glorioso, Paolo, et al.
Publicado: (2024) -
Zyda: A 1.3T Dataset for Open Language Modeling
por: Tokpanov, Yury, et al.
Publicado: (2024) -
BlackMamba: Mixture of Experts for State-Space Models
por: Anthony, Quentin, et al.
Publicado: (2024) -
Zyda-2: a 5 Trillion Token High-Quality Dataset
por: Tokpanov, Yury, et al.
Publicado: (2024) -
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
por: Shyam, Vasudev, et al.
Publicado: (2024)