Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Jingze, Wu, Bingheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910755286155264
author Shi, Jingze
Wu, Bingheng
author_facet Shi, Jingze
Wu, Bingheng
contents In order to make the foundation model more efficient and effective, our idea is combining sequence transformation and state transformation. First, we prove the availability of rotary position embedding in the state space duality algorithm, which reduces the perplexity of the hybrid quadratic causal self-attention and state space duality by more than 4%, to ensure that the combining sequence transformation unifies position encoding. Second, we propose dynamic mask attention, which maintains 100% accuracy in the more challenging multi-query associative recall task, improving by more than 150% compared to quadratic causal self-attention and state space duality, to ensure that the combining sequence transformation selectively filters relevant information. Third, we design cross domain mixture of experts, which makes the computational speed of expert retrieval with more than 1024 experts 8 to 10 times faster than the mixture of experts, to ensure that the combining state transformation quickly retrieval mixture. Finally, we summarize these matrix algorithms that can form the foundation model: Wonderful Matrices, which can be a competitor to popular model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11834
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture
Shi, Jingze
Wu, Bingheng
Machine Learning
Artificial Intelligence
Computation and Language
In order to make the foundation model more efficient and effective, our idea is combining sequence transformation and state transformation. First, we prove the availability of rotary position embedding in the state space duality algorithm, which reduces the perplexity of the hybrid quadratic causal self-attention and state space duality by more than 4%, to ensure that the combining sequence transformation unifies position encoding. Second, we propose dynamic mask attention, which maintains 100% accuracy in the more challenging multi-query associative recall task, improving by more than 150% compared to quadratic causal self-attention and state space duality, to ensure that the combining sequence transformation selectively filters relevant information. Third, we design cross domain mixture of experts, which makes the computational speed of expert retrieval with more than 1024 experts 8 to 10 times faster than the mixture of experts, to ensure that the combining state transformation quickly retrieval mixture. Finally, we summarize these matrix algorithms that can form the foundation model: Wonderful Matrices, which can be a competitor to popular model architectures.
title Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.11834