Shared DIFF Transformer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cang, Yueyang, Liu, Yuhang, Zhang, Xiaoteng, Shi, Li, Que, Wenge
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914202179862528
author Cang, Yueyang
Liu, Yuhang
Zhang, Xiaoteng
Shi, Li
Que, Wenge
author_facet Cang, Yueyang
Liu, Yuhang
Zhang, Xiaoteng
Shi, Li
Que, Wenge
contents DIFF Transformer improves attention allocation by enhancing focus on relevant context while suppressing noise. It introduces a differential attention mechanism that calculates the difference between two independently generated attention distributions, effectively reducing noise and promoting sparse attention patterns. However, the independent signal generation in DIFF Transformer results in parameter redundancy and suboptimal utilization of information. In this work, we propose Shared DIFF Transformer, which draws on the idea of a differential amplifier by introducing a shared base matrix to model global patterns and incorporating low-rank updates to enhance task-specific flexibility. This design significantly reduces parameter redundancy, improves efficiency, and retains strong noise suppression capabilities. Experimental results show that, compared to DIFF Transformer, our method achieves better performance in tasks such as long-sequence modeling, key information retrieval, and in-context learning. Our work provides a novel and efficient approach to optimizing differential attention mechanisms and advancing robust Transformer architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Shared DIFF Transformer
Cang, Yueyang
Liu, Yuhang
Zhang, Xiaoteng
Shi, Li
Que, Wenge
Machine Learning
DIFF Transformer improves attention allocation by enhancing focus on relevant context while suppressing noise. It introduces a differential attention mechanism that calculates the difference between two independently generated attention distributions, effectively reducing noise and promoting sparse attention patterns. However, the independent signal generation in DIFF Transformer results in parameter redundancy and suboptimal utilization of information. In this work, we propose Shared DIFF Transformer, which draws on the idea of a differential amplifier by introducing a shared base matrix to model global patterns and incorporating low-rank updates to enhance task-specific flexibility. This design significantly reduces parameter redundancy, improves efficiency, and retains strong noise suppression capabilities. Experimental results show that, compared to DIFF Transformer, our method achieves better performance in tasks such as long-sequence modeling, key information retrieval, and in-context learning. Our work provides a novel and efficient approach to optimizing differential attention mechanisms and advancing robust Transformer architectures.
title Shared DIFF Transformer
topic Machine Learning
url https://arxiv.org/abs/2501.17900