EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiyu, Peng, Shuang, Luo, Daxiong, Yang, Fan, Wu, Renshou, Li, Fangyuan, Chen, Xiaoxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913763735633920
author Chen, Jiyu
Peng, Shuang
Luo, Daxiong
Yang, Fan
Wu, Renshou
Li, Fangyuan
Chen, Xiaoxin
author_facet Chen, Jiyu
Peng, Shuang
Luo, Daxiong
Yang, Fan
Wu, Renshou
Li, Fangyuan
Chen, Xiaoxin
contents Transformer-based large language models (LLMs) encounter challenges in processing long sequences on edge devices due to the quadratic complexity of attention mechanisms and growing memory demands from Key-Value (KV) cache. Existing KV cache optimizations struggle with irreversible token eviction in long-output tasks, while alternative sequence modeling architectures prove costly to adopt within established Transformer infrastructure. We present EdgeInfinite, a memory-efficient solution for infinite contexts that integrates compressed memory into Transformer-based LLMs through a trainable memory-gating module. This approach maintains full compatibility with standard Transformer architectures, requiring fine-tuning only a small part of parameters, and enables selective activation of the memory-gating module for long and short context task routing. The experimental result shows that EdgeInfinite achieves comparable performance to baseline Transformer-based LLM on long context benchmarks while optimizing memory consumption and time to first token.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices
Chen, Jiyu
Peng, Shuang
Luo, Daxiong
Yang, Fan
Wu, Renshou
Li, Fangyuan
Chen, Xiaoxin
Computation and Language
Transformer-based large language models (LLMs) encounter challenges in processing long sequences on edge devices due to the quadratic complexity of attention mechanisms and growing memory demands from Key-Value (KV) cache. Existing KV cache optimizations struggle with irreversible token eviction in long-output tasks, while alternative sequence modeling architectures prove costly to adopt within established Transformer infrastructure. We present EdgeInfinite, a memory-efficient solution for infinite contexts that integrates compressed memory into Transformer-based LLMs through a trainable memory-gating module. This approach maintains full compatibility with standard Transformer architectures, requiring fine-tuning only a small part of parameters, and enables selective activation of the memory-gating module for long and short context task routing. The experimental result shows that EdgeInfinite achieves comparable performance to baseline Transformer-based LLM on long context benchmarks while optimizing memory consumption and time to first token.
title EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices
topic Computation and Language
url https://arxiv.org/abs/2503.22196