Supernova: Achieving More with Less in Transformer Architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tanase, Andrei-Valentin, Pelican, Elena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912496068067328
author Tanase, Andrei-Valentin
Pelican, Elena
author_facet Tanase, Andrei-Valentin
Pelican, Elena
contents We present Supernova, a 650M-parameter decoder-only transformer that demonstrates how careful architectural design and tokenization innovation can achieve the performance of larger models while maintaining computational efficiency. Our architecture combines Rotary Positional Embeddings (RoPE), Grouped Query Attention (GQA) with a 3:1 compression ratio, RMSNorm for computational efficiency, and SwiGLU activation functions. A critical innovation is our custom 128,000-vocabulary byte-level BPE tokenizer, which achieves state-of-the-art compression performance. Through detailed analysis, we show that Supernova achieves 90% of the performance of 1B-parameter models while using 35% fewer parameters and requiring only 100B training tokens--an order of magnitude less than competing models. Our findings challenge the prevailing scaling paradigm, demonstrating that architectural efficiency and tokenization quality can compensate for reduced parameter counts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Supernova: Achieving More with Less in Transformer Architectures
Tanase, Andrei-Valentin
Pelican, Elena
Computation and Language
Artificial Intelligence
Machine Learning
We present Supernova, a 650M-parameter decoder-only transformer that demonstrates how careful architectural design and tokenization innovation can achieve the performance of larger models while maintaining computational efficiency. Our architecture combines Rotary Positional Embeddings (RoPE), Grouped Query Attention (GQA) with a 3:1 compression ratio, RMSNorm for computational efficiency, and SwiGLU activation functions. A critical innovation is our custom 128,000-vocabulary byte-level BPE tokenizer, which achieves state-of-the-art compression performance. Through detailed analysis, we show that Supernova achieves 90% of the performance of 1B-parameter models while using 35% fewer parameters and requiring only 100B training tokens--an order of magnitude less than competing models. Our findings challenge the prevailing scaling paradigm, demonstrating that architectural efficiency and tokenization quality can compensate for reduced parameter counts.
title Supernova: Achieving More with Less in Transformer Architectures
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.15773