Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Woomin, Jayanthi, Sai Muralidhar, Ronanki, Srikanth, Sathyendra, Kanthashree Mysore, Shin, Jinwoo, Galstyan, Aram, Katiyar, Shubham, Bodapati, Sravan Babu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908658392104960
author Song, Woomin
Jayanthi, Sai Muralidhar
Ronanki, Srikanth
Sathyendra, Kanthashree Mysore
Shin, Jinwoo
Galstyan, Aram
Katiyar, Shubham
Bodapati, Sravan Babu
author_facet Song, Woomin
Jayanthi, Sai Muralidhar
Ronanki, Srikanth
Sathyendra, Kanthashree Mysore
Shin, Jinwoo
Galstyan, Aram
Katiyar, Shubham
Bodapati, Sravan Babu
contents As large language models increasingly gain popularity in real-world applications, processing extremely long contexts, often exceeding the model's pre-trained context limits, has emerged as a critical challenge. While existing approaches to efficient long-context processing show promise, recurrent compression-based methods struggle with information preservation, whereas random access approaches require substantial memory resources. We introduce REFORM, a novel inference framework that efficiently handles long contexts through a two-phase approach. First, it incrementally processes input chunks while maintaining a compressed KV cache, constructs cross-layer context embeddings, and utilizes early exit strategy for improved efficiency. Second, it identifies and gathers essential tokens via similarity matching and selectively recomputes the KV cache. Compared to baselines, REFORM achieves over 52% and 34% performance gains on RULER and BABILong respectively at 1M context length. It also outperforms baselines on Infinite-Bench, RepoEval, and MM-NIAH, demonstrating flexibility across diverse tasks and domains. Additionally, REFORM reduces inference time by 30% and peak memory usage by 5%, achieving both efficiency and superior performance.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
Song, Woomin
Jayanthi, Sai Muralidhar
Ronanki, Srikanth
Sathyendra, Kanthashree Mysore
Shin, Jinwoo
Galstyan, Aram
Katiyar, Shubham
Bodapati, Sravan Babu
Computation and Language
Machine Learning
As large language models increasingly gain popularity in real-world applications, processing extremely long contexts, often exceeding the model's pre-trained context limits, has emerged as a critical challenge. While existing approaches to efficient long-context processing show promise, recurrent compression-based methods struggle with information preservation, whereas random access approaches require substantial memory resources. We introduce REFORM, a novel inference framework that efficiently handles long contexts through a two-phase approach. First, it incrementally processes input chunks while maintaining a compressed KV cache, constructs cross-layer context embeddings, and utilizes early exit strategy for improved efficiency. Second, it identifies and gathers essential tokens via similarity matching and selectively recomputes the KV cache. Compared to baselines, REFORM achieves over 52% and 34% performance gains on RULER and BABILong respectively at 1M context length. It also outperforms baselines on Infinite-Bench, RepoEval, and MM-NIAH, demonstrating flexibility across diverse tasks and domains. Additionally, REFORM reduces inference time by 30% and peak memory usage by 5%, achieving both efficiency and superior performance.
title Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.01215