DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Picón, Ginés Carreto, Zhou, Peng Yuan, Zhang, Qi, Iosifidis, Alexandros
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917374150574080
author Picón, Ginés Carreto
Zhou, Peng Yuan
Zhang, Qi
Iosifidis, Alexandros
author_facet Picón, Ginés Carreto
Zhou, Peng Yuan
Zhang, Qi
Iosifidis, Alexandros
contents Transformer-based models have dramatically increased their size and parameter count to tackle increasingly complex tasks. At the same time, there is a growing demand for high performance, low-latency inference on devices with limited resources. In particular, stream data inference is typically performed over a sliding temporal window, leading to highly redundant computations. While the recent Continual Transformers started addressing this issue, they can be effectively used only in shallow models, which limits their scope and generalization power. In this paper, we propose the Deep Continual Transformer (DeepCoT), a redundancy-free encoder attention mechanism that can be applied over existing deep encoder architectures with minimal changes. In our experiments over audio, video, and text streams, we show that DeepCoTs retain comparative performance to their non-continual baselines while offering a linear computational cost for all Transformer layers, which reduces up to two orders of magnitude in the running time compared to previous efficient models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17693
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
Picón, Ginés Carreto
Zhou, Peng Yuan
Zhang, Qi
Iosifidis, Alexandros
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Transformer-based models have dramatically increased their size and parameter count to tackle increasingly complex tasks. At the same time, there is a growing demand for high performance, low-latency inference on devices with limited resources. In particular, stream data inference is typically performed over a sliding temporal window, leading to highly redundant computations. While the recent Continual Transformers started addressing this issue, they can be effectively used only in shallow models, which limits their scope and generalization power. In this paper, we propose the Deep Continual Transformer (DeepCoT), a redundancy-free encoder attention mechanism that can be applied over existing deep encoder architectures with minimal changes. In our experiments over audio, video, and text streams, we show that DeepCoTs retain comparative performance to their non-continual baselines while offering a linear computational cost for all Transformer layers, which reduces up to two orders of magnitude in the running time compared to previous efficient models.
title DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17693