DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917374150574080 |
|---|---|
| author | Picón, Ginés Carreto Zhou, Peng Yuan Zhang, Qi Iosifidis, Alexandros |
| author_facet | Picón, Ginés Carreto Zhou, Peng Yuan Zhang, Qi Iosifidis, Alexandros |
| contents | Transformer-based models have dramatically increased their size and parameter count to tackle increasingly complex tasks. At the same time, there is a growing demand for high performance, low-latency inference on devices with limited resources. In particular, stream data inference is typically performed over a sliding temporal window, leading to highly redundant computations. While the recent Continual Transformers started addressing this issue, they can be effectively used only in shallow models, which limits their scope and generalization power. In this paper, we propose the Deep Continual Transformer (DeepCoT), a redundancy-free encoder attention mechanism that can be applied over existing deep encoder architectures with minimal changes. In our experiments over audio, video, and text streams, we show that DeepCoTs retain comparative performance to their non-continual baselines while offering a linear computational cost for all Transformer layers, which reduces up to two orders of magnitude in the running time compared to previous efficient models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_17693 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams Picón, Ginés Carreto Zhou, Peng Yuan Zhang, Qi Iosifidis, Alexandros Machine Learning Computation and Language Computer Vision and Pattern Recognition Transformer-based models have dramatically increased their size and parameter count to tackle increasingly complex tasks. At the same time, there is a growing demand for high performance, low-latency inference on devices with limited resources. In particular, stream data inference is typically performed over a sliding temporal window, leading to highly redundant computations. While the recent Continual Transformers started addressing this issue, they can be effectively used only in shallow models, which limits their scope and generalization power. In this paper, we propose the Deep Continual Transformer (DeepCoT), a redundancy-free encoder attention mechanism that can be applied over existing deep encoder architectures with minimal changes. In our experiments over audio, video, and text streams, we show that DeepCoTs retain comparative performance to their non-continual baselines while offering a linear computational cost for all Transformer layers, which reduces up to two orders of magnitude in the running time compared to previous efficient models. |
| title | DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams |
| topic | Machine Learning Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.17693 |