Saved in:
| Main Authors: | Glorioso, Paolo, Anthony, Quentin, Tokpanov, Yury, Golubeva, Anna, Shyam, Vasudev, Whittington, James, Pilault, Jonathan, Millidge, Beren |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.15242 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024)
by: Glorioso, Paolo, et al.
Published: (2024)
Zyda: A 1.3T Dataset for Open Language Modeling
by: Tokpanov, Yury, et al.
Published: (2024)
by: Tokpanov, Yury, et al.
Published: (2024)
BlackMamba: Mixture of Experts for State-Space Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Zyda-2: a 5 Trillion Token High-Quality Dataset
by: Tokpanov, Yury, et al.
Published: (2024)
by: Tokpanov, Yury, et al.
Published: (2024)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
by: Shyam, Vasudev, et al.
Published: (2024)
by: Shyam, Vasudev, et al.
Published: (2024)
ZAYA1-8B Technical Report
by: Washbourne, Robert, et al.
Published: (2026)
by: Washbourne, Robert, et al.
Published: (2026)
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
by: Figliolia, Tomas, et al.
Published: (2025)
by: Figliolia, Tomas, et al.
Published: (2025)
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
by: Anthony, Quentin, et al.
Published: (2025)
by: Anthony, Quentin, et al.
Published: (2025)
ZAYA1-VL-8B Technical Report
by: Shapourian, Hassan, et al.
Published: (2026)
by: Shapourian, Hassan, et al.
Published: (2026)
Hybrid Associative Memories
by: Lufkin, Leon, et al.
Published: (2026)
by: Lufkin, Leon, et al.
Published: (2026)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
by: Alonso, Nick, et al.
Published: (2024)
by: Alonso, Nick, et al.
Published: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
by: Shyam, Vasu, et al.
Published: (2026)
by: Shyam, Vasu, et al.
Published: (2026)
ZUNA: Flexible EEG Superresolution with Position-Aware Diffusion Autoencoders
by: Warner, Christopher, et al.
Published: (2026)
by: Warner, Christopher, et al.
Published: (2026)
Generalising E-prop to Deep Networks
by: Millidge, Beren
Published: (2025)
by: Millidge, Beren
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
PLaMo 2 Technical Report
by: Networks, Preferred, et al.
Published: (2025)
by: Networks, Preferred, et al.
Published: (2025)
Nemotron-4 15B Technical Report
by: Parmar, Jupinder, et al.
Published: (2024)
by: Parmar, Jupinder, et al.
Published: (2024)
Yi-Lightning Technical Report
by: Wake, Alan, et al.
Published: (2024)
by: Wake, Alan, et al.
Published: (2024)
PCMind-2.1-Kaiyuan-2B Technical Report
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Nemotron-4 340B Technical Report
by: Nvidia, et al.
Published: (2024)
by: Nvidia, et al.
Published: (2024)
Trillion 7B Technical Report
by: Han, Sungjun, et al.
Published: (2025)
by: Han, Sungjun, et al.
Published: (2025)
Skywork Open Reasoner 1 Technical Report
by: He, Jujie, et al.
Published: (2025)
by: He, Jujie, et al.
Published: (2025)
EuroLLM-22B: Technical Report
by: Ramos, Miguel Moura, et al.
Published: (2026)
by: Ramos, Miguel Moura, et al.
Published: (2026)
EuroLLM-9B: Technical Report
by: Martins, Pedro Henrique, et al.
Published: (2025)
by: Martins, Pedro Henrique, et al.
Published: (2025)
Exploring Action-Centric Representations Through the Lens of Rate-Distortion Theory
by: Varona, Miguel de Llanza, et al.
Published: (2024)
by: Varona, Miguel de Llanza, et al.
Published: (2024)
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG
by: Alonso, Nicholas, et al.
Published: (2024)
by: Alonso, Nicholas, et al.
Published: (2024)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
by: Rohekar, Raanan Y., et al.
Published: (2024)
by: Rohekar, Raanan Y., et al.
Published: (2024)
Technical Report: Small Language Model for Japanese Clinical and Medicine
by: Watanabe, Shogo
Published: (2024)
by: Watanabe, Shogo
Published: (2024)
Ovis2.5 Technical Report
by: Lu, Shiyin, et al.
Published: (2025)
by: Lu, Shiyin, et al.
Published: (2025)
Latxa: An Open Language Model and Evaluation Suite for Basque
by: Etxaniz, Julen, et al.
Published: (2024)
by: Etxaniz, Julen, et al.
Published: (2024)
Kwai Summary Attention Technical Report
by: Chu, Chenglong, et al.
Published: (2026)
by: Chu, Chenglong, et al.
Published: (2026)
TTS-1 Technical Report
by: Atamanenko, Oleg, et al.
Published: (2025)
by: Atamanenko, Oleg, et al.
Published: (2025)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
by: Yang, An, et al.
Published: (2024)
by: Yang, An, et al.
Published: (2024)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)
by: Gohil, Vasudev
Published: (2025)
Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
by: Dade, Nii Osae Osae, et al.
Published: (2025)
by: Dade, Nii Osae Osae, et al.
Published: (2025)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Probing Semantic Routing in Large Mixture-of-Expert Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding Networks
by: Salvatori, Tommaso, et al.
Published: (2022)
by: Salvatori, Tommaso, et al.
Published: (2022)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
by: Iyer, Vivek, et al.
Published: (2025)
by: Iyer, Vivek, et al.
Published: (2025)
Similar Items
-
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024) -
Zyda: A 1.3T Dataset for Open Language Modeling
by: Tokpanov, Yury, et al.
Published: (2024) -
BlackMamba: Mixture of Experts for State-Space Models
by: Anthony, Quentin, et al.
Published: (2024) -
Zyda-2: a 5 Trillion Token High-Quality Dataset
by: Tokpanov, Yury, et al.
Published: (2024) -
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
by: Shyam, Vasudev, et al.
Published: (2024)