No Mean Feat: Simple, Strong Baselines for Context Compression

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Feldman, Yair, Artzi, Yoav
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915999493652480
author Feldman, Yair
Artzi, Yoav
author_facet Feldman, Yair
Artzi, Yoav
contents Context compression reduces Transformer inference costs by replacing lengthy inputs with shorter pre-computed representations. It carries significant benefits for retrieval-augmented generation (RAG) and has attracted growing research attention. However, progress remains difficult to measure due to inconsistent evaluations and baselines. We design a standard, easy-to-reproduce evaluation suite for context compression, BenchPress, along with simple, high-performance baselines for English reading comprehension. BenchPress supports benchmarking across model scales, datasets, compression ratios, and short ($<$1K tokens) to mid-range ($<$8K tokens) contexts. While the suite is applicable to any compression paradigm, our baselines target soft context compression. We establish two simple baselines that strongly outperform the widely used causal compression-token approach: mean pooling and a bidirectional compression-token variant. Our results show the benefit of bidirectional attention when computing compressed representations, and that simple pooling is an expressive compression operator.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle No Mean Feat: Simple, Strong Baselines for Context Compression
Feldman, Yair
Artzi, Yoav
Computation and Language
Artificial Intelligence
Machine Learning
Context compression reduces Transformer inference costs by replacing lengthy inputs with shorter pre-computed representations. It carries significant benefits for retrieval-augmented generation (RAG) and has attracted growing research attention. However, progress remains difficult to measure due to inconsistent evaluations and baselines. We design a standard, easy-to-reproduce evaluation suite for context compression, BenchPress, along with simple, high-performance baselines for English reading comprehension. BenchPress supports benchmarking across model scales, datasets, compression ratios, and short ($<$1K tokens) to mid-range ($<$8K tokens) contexts. While the suite is applicable to any compression paradigm, our baselines target soft context compression. We establish two simple baselines that strongly outperform the widely used causal compression-token approach: mean pooling and a bidirectional compression-token variant. Our results show the benefit of bidirectional attention when computing compressed representations, and that simple pooling is an expressive compression operator.
title No Mean Feat: Simple, Strong Baselines for Context Compression
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.20797