EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913763735633920 |
|---|---|
| author | Chen, Jiyu Peng, Shuang Luo, Daxiong Yang, Fan Wu, Renshou Li, Fangyuan Chen, Xiaoxin |
| author_facet | Chen, Jiyu Peng, Shuang Luo, Daxiong Yang, Fan Wu, Renshou Li, Fangyuan Chen, Xiaoxin |
| contents | Transformer-based large language models (LLMs) encounter challenges in processing long sequences on edge devices due to the quadratic complexity of attention mechanisms and growing memory demands from Key-Value (KV) cache. Existing KV cache optimizations struggle with irreversible token eviction in long-output tasks, while alternative sequence modeling architectures prove costly to adopt within established Transformer infrastructure. We present EdgeInfinite, a memory-efficient solution for infinite contexts that integrates compressed memory into Transformer-based LLMs through a trainable memory-gating module. This approach maintains full compatibility with standard Transformer architectures, requiring fine-tuning only a small part of parameters, and enables selective activation of the memory-gating module for long and short context task routing. The experimental result shows that EdgeInfinite achieves comparable performance to baseline Transformer-based LLM on long context benchmarks while optimizing memory consumption and time to first token. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_22196 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices Chen, Jiyu Peng, Shuang Luo, Daxiong Yang, Fan Wu, Renshou Li, Fangyuan Chen, Xiaoxin Computation and Language Transformer-based large language models (LLMs) encounter challenges in processing long sequences on edge devices due to the quadratic complexity of attention mechanisms and growing memory demands from Key-Value (KV) cache. Existing KV cache optimizations struggle with irreversible token eviction in long-output tasks, while alternative sequence modeling architectures prove costly to adopt within established Transformer infrastructure. We present EdgeInfinite, a memory-efficient solution for infinite contexts that integrates compressed memory into Transformer-based LLMs through a trainable memory-gating module. This approach maintains full compatibility with standard Transformer architectures, requiring fine-tuning only a small part of parameters, and enables selective activation of the memory-gating module for long and short context task routing. The experimental result shows that EdgeInfinite achieves comparable performance to baseline Transformer-based LLM on long context benchmarks while optimizing memory consumption and time to first token. |
| title | EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2503.22196 |