Superposition in Transformers: A Novel Way of Building Mixture of Experts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chaliah, Ayoub Ben, Dellagi, Hela
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913637456674816
author Chaliah, Ayoub Ben
Dellagi, Hela
author_facet Chaliah, Ayoub Ben
Dellagi, Hela
contents Catastrophic forgetting remains a major challenge when adapting large language models (LLMs) to new tasks or domains. Conventional fine-tuning often overwrites existing knowledge, causing performance degradation on original tasks. We introduce Superposition in Transformers, a novel architecture that leverages autoencoders to superimpose the hidden representations of a base model and a fine-tuned model within a shared parameter space. By using B-spline-based blending coefficients and autoencoders that adaptively reconstruct hidden states based on the input data distribution, our method effectively mitigates catastrophic forgetting and enables a new paradigm of "in-model" superposition. This approach preserves original model capabilities while allowing compact domain-specific expertise to be added, and it supports dynamic switching between model states during inference.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Superposition in Transformers: A Novel Way of Building Mixture of Experts
Chaliah, Ayoub Ben
Dellagi, Hela
Computation and Language
Artificial Intelligence
Catastrophic forgetting remains a major challenge when adapting large language models (LLMs) to new tasks or domains. Conventional fine-tuning often overwrites existing knowledge, causing performance degradation on original tasks. We introduce Superposition in Transformers, a novel architecture that leverages autoencoders to superimpose the hidden representations of a base model and a fine-tuned model within a shared parameter space. By using B-spline-based blending coefficients and autoencoders that adaptively reconstruct hidden states based on the input data distribution, our method effectively mitigates catastrophic forgetting and enables a new paradigm of "in-model" superposition. This approach preserves original model capabilities while allowing compact domain-specific expertise to be added, and it supports dynamic switching between model states during inference.
title Superposition in Transformers: A Novel Way of Building Mixture of Experts
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.00530